Adapting the Tesseract open source OCR engine for multilingual OCR

80Citations
Citations of this article
194Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We describe efforts to adapt the Tesseract open source OCR engine for multiple scripts and languages. Effort has been concentrated on enabling generic multi-lingual operation such that negligible customization is required for a new language beyond providing a corpus of text. Although change was required to various modules, including physical layout analysis, and linguistic post-processing, no change was required to the character classifier beyond changing a few limits. The Tesseract classifier has adapted easily to Simplified Chinese. Test results on English, a mixture of European languages, and Russian, taken from a random sample of books, show a reasonably consistent word error rate between 3.72% and 5.78%, and Simplified Chinese has a character error rate of only 3.77%. Copyright © 2009 ACM.

Author supplied keywords

Cite

CITATION STYLE

APA

Smith, R., Antonova, D., & Lee, D. S. (2009). Adapting the Tesseract open source OCR engine for multilingual OCR. In ACM International Conference Proceeding Series. Association for Computing Machinery. https://doi.org/10.1145/1577802.1577804

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free