Abstract
We describe efforts to adapt the Tesseract open source OCR engine for multiple scripts and languages. Effort has been concentrated on enabling generic multi-lingual operation such that negligible customization is required for a new language beyond providing a corpus of text. Although change was required to various modules, including physical layout analysis, and linguistic post-processing, no change was required to the character classifier beyond changing a few limits. The Tesseract classifier has adapted easily to Simplified Chinese. Test results on English, a mixture of European languages, and Russian, taken from a random sample of books, show a reasonably consistent word error rate between 3.72% and 5.78%, and Simplified Chinese has a character error rate of only 3.77%. Copyright © 2009 ACM.
Author supplied keywords
Cite
CITATION STYLE
Smith, R., Antonova, D., & Lee, D. S. (2009). Adapting the Tesseract open source OCR engine for multilingual OCR. In ACM International Conference Proceeding Series. Association for Computing Machinery. https://doi.org/10.1145/1577802.1577804
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.