Computer Science > Computer Vision and Pattern Recognition

arXiv:1912.07025 (cs)

[Submitted on 15 Dec 2019]

Title:Indiscapes: Instance Segmentation Networks for Layout Parsing of Historical Indic Manuscripts

Authors:Abhishek Prusty, Sowmya Aitha, Abhishek Trivedi, Ravi Kiran Sarvadevabhatla

View PDF

Abstract:Historical palm-leaf manuscript and early paper documents from Indian subcontinent form an important part of the world's literary and cultural heritage. Despite their importance, large-scale annotated Indic manuscript image datasets do not exist. To address this deficiency, we introduce Indiscapes, the first ever dataset with multi-regional layout annotations for historical Indic manuscripts. To address the challenge of large diversity in scripts and presence of dense, irregular layout elements (e.g. text lines, pictures, multiple documents per image), we adapt a Fully Convolutional Deep Neural Network architecture for fully automatic, instance-level spatial layout parsing of manuscript images. We demonstrate the effectiveness of proposed architecture on images from the Indiscapes dataset. For annotation flexibility and keeping the non-technical nature of domain experts in mind, we also contribute a custom, web-based GUI annotation tool and a dashboard-style analytics portal. Overall, our contributions set the stage for enabling downstream applications such as OCR and word-spotting in historical Indic manuscripts at scale.

Comments:	Oral presentation at International Conference on Document Analysis and Recognition (ICDAR) - 2019. For dataset, pre-trained networks and additional details, visit project page at this http URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as:	arXiv:1912.07025 [cs.CV]
	(or arXiv:1912.07025v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1912.07025

Submission history

From: Ravi Kiran Sarvadevabhatla [view email]
[v1] Sun, 15 Dec 2019 11:42:27 UTC (9,565 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Indiscapes: Instance Segmentation Networks for Layout Parsing of Historical Indic Manuscripts

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Indiscapes: Instance Segmentation Networks for Layout Parsing of Historical Indic Manuscripts

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators