Abstract
Optical Character Recognition (OCR) is still hard in uncontrolled settings, where poor lighting—like low light, glare, shadows, and mixed illumination—reduces text visibility and recognition accuracy. To tackle this, a hybrid OCR pipeline is proposed that combines optimized pre-processing, deep-learning recognition, and semantic post-processing with a Large Language Model (LLM). Contrast Limited Adaptive Histogram Equalization (CLAHE), Gamma Correction, and Maximally Stable Extremal Region (MSER) detection, with MSER parameters are tuned by the Bacterial Foraging Optimization Algorithm (BFOA) for more reliable text localization. The pre-processed images are fed to a Convolutional Recurrent Neural Network (CRNN) with Connectionist Temporal Classification (CTC) decoding for recognition. Finally, a GPT-4-based LLM fixes semantic and contextual errors in the extracted text. We tested our method on 10,000 real-world scene text images captured under systematically varied lighting. Traditional OCR pipelines reached only moderate robustness (about 78–92% similarity ratio), and deep-learning OCR improved recognition to around 93%. Our hybrid approach reached a 97.0% similarity ratio, delivering both syntactic accuracy and semantic correctness. Ablation studies show each stage is important, and runtime analysis highlights the trade-off between accuracy and computational cost. Against recent state-of-the-art OCR systems, proposed framework consistently performs better under extreme illumination, showing reliability and scalability for real-world use.
Author supplied keywords
Cite
CITATION STYLE
Rakesh, T. M., Girisha, G. S., & Renukadevi, M. N. (2025). Hybrid OCR with LLM -Enhanced Post Processing for Robust Text Recognition for Extreme Illumination Condition. International Journal of Intelligent Engineering and Systems, 18(11), 1049–1064. https://doi.org/10.22266/ijies2025.1231.65
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.