Abstract
In this paper, we describe our system HASP-2015 (Hybrid Arabic Spelling and Punctuation Corrector) in which we introduce significant improvements over our previous version HASP-2014 and with which we participated in the QALB- 2015 Second Shared Task on Arabic Error Correction. Our system utilizes probabilistic information on errors and their possible corrections in the training data and combine that with an open-source reference dictionary (or word list) for detecting errors and generating and filtering candidates. We enhance our system further by allowing it to generate candidates for common semantic and grammatical errors. Eventually, an n-gram language model is used for selecting best candidates. We use a CRF (Conditional Random Fields) classifier for correcting punctuation errors in a two-pass process where first the system learns punctuation placement, and then it learns to identify punctuation types.
Cite
CITATION STYLE
Attia, M., Al-Badrashiny, M., & Diab, M. (2015). Gwu-hasp-2015@qalb-2015 shared task: Priming spelling candidates with probability1. In 2nd Workshop on Arabic Natural Language Processing, ANLP 2015 - held at 53rd Annual Meeting of the Association for Computational Linguistics, ACL 2015 - Proceedings (pp. 138–143). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-3216
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.