Single-Line Text Detection in Multi-Line Text with Narrow Spacing for Line-Based Character Recognition

3Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

Text detection is a crucial pre-processing step in optical character recognition (OCR) for the accurate recognition of text, including both fonts and handwritten characters, in documents. While current deep learning-based text detection tools can detect text regions with high accuracy, they often treat multiple lines of text as a single region. To perform line-based character recognition, it is necessary to divide the text into individual lines, which requires a line detection technique. This paper focuses on the development of a new approach to single-line detection in OCR that is based on the existing Character Region Awareness For Text detection (CRAFT) model and incorporates a deep neural network specialized in line segmentation. However, this new method may still detect multiple lines as a single text region when multi-line text with narrow spacing is present. To address this, we also introduce a post-processing algorithm to detect single text regions using the output of the single-line segmentation. Our proposed method successfully detects single lines, even in multi-line text with narrow line spacing, and hence improves the accuracy of OCR.

Cite

CITATION STYLE

APA

Leow, C. S., Yajima, H., Kitagawa, T., & Nishizaki, H. (2023). Single-Line Text Detection in Multi-Line Text with Narrow Spacing for Line-Based Character Recognition. IEICE Transactions on Information and Systems, E106.D(12), 2097–2106. https://doi.org/10.1587/transinf.2023EDP7070

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free