Automatic Hindi OCR Error Correction Using MLM-BERT

3Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

Optical Character Recognition (OCR) systems find it challenging to generate accurate text for highly inflectional Indic languages such as Hindi. Inflectional languages possess an extensive vocabulary. Words in these languages can assume different forms based on factors like gender, meaning, or other contextual cues. To enhance the accuracy of OCR and correct the errors resulting from the inflectional nature of language, it is crucial to perform post-processing on output of the OCR. This work focuses on correcting errors in the OCR output specifically for the Hindi language. To overcome existing challenges, an error correction model has been proposed in this work that uses the Masked-Language Modeling with BERT (MLM-BERT). It utilizes the context to provide accurate word suggestions for the incorrect word or masked word. The proposed model has been tested using the Hindi OCR test dataset from IIITH. It achieved an improvement of 3.58% word accuracy over the baseline OCR word accuracy, which demonstrates its effectiveness in enhancing the accuracy of the OCR output text.

Cite

CITATION STYLE

APA

Kundaikar, T., Fadte, S., Karmali, R., & Pawar, J. D. (2024). Automatic Hindi OCR Error Correction Using MLM-BERT. Ingenierie Des Systemes d’Information, 29(2), 619–626. https://doi.org/10.18280/isi.290223

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free