Text region extraction from quality degraded document images

S. Abirami; D. Manjula

Conference ProceedingsOPEN ACCESS

Text region extraction from quality degraded document images

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2007) 4815 LNCS 519-527

DOI: 10.1007/978-3-540-77046-6_64

0Citations

3Readers

Abstract

In this paper we present a well designed method that makes use of edge information to extract textual blocks from gray scale document images. It aims at detecting textual regions on heavy noise infected newspaper images and separate them from graphical regions. The algorithm traces the feature points in different entities and then groups those edge points of textual regions. Finally feature based connected component merging was introduced to gather homogeneous textual regions together within the scope of its bounding rectangles. The proposed method can be used to locate text in-group of newspaper images with multiple page layouts. Initial results are encouraging, then they are experimented with considerable number of newspaper images with different layout structures and promising results were obtained. This finds its major application in digital libraries for OCR where information can be of different quality depending on the age of the scanned paper. © Springer-Verlag Berlin Heidelberg 2007.

Author supplied keywords

Cite

CITATION STYLE

APA

Abirami, S., & Manjula, D. (2007). Text region extraction from quality degraded document images. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 4815 LNCS, pp. 519–527). Springer Verlag. https://doi.org/10.1007/978-3-540-77046-6_64

Text region extraction from quality degraded document images

Abstract

Author supplied keywords

Cite

Register to see more suggestions