Abstract
The recognition of Arabic diacritic symbols is essential for accurate comprehension of Arabic texts, yet most existing Optical Character Recognition (OCR) systems lack this capability. This paper proposes a novel approach for Arabic diacritic recognition using a domain adaptation method applied to the AlexNet deep convolutional neural network architecture. A custom dataset of synthetic Arabic diacritic images was generated, incorporating variations in font styles, sizes, slant angles, and perspective distortions. The model was trained using this dataset, and a domain adaptation technique was employed to improve its generalization across diverse diacritic styles and handwriting variations. We evaluated the proposed model on the testing dataset consisting of a variety of fonts and distortions; the result was an overall accuracy of 97%, indicating that the model can achieve a high accuracy for most diacritic categories, with only a few misclassifications.
Cite
CITATION STYLE
Alalawi, Y., Chandler, D. M., & Caluya, N. R. (2023). A CNN-Based Arabic Diacritic Symbol Recognition System Using Domain Adaptation. In ACM International Conference Proceeding Series (pp. 23–32). Association for Computing Machinery. https://doi.org/10.1145/3626641.3627212
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.