Abstract
We present MICHAEL, a lightweight method developed for the MADAR shared task on travel domain Dialect Identification (DID). It uses character-level features and perform classification without any pre-processing. Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier. MICHAEL achieved an official score (accuracy) of 53.25% with 1 ≤ N ≤ 3 but showed a much better result with character 4-grams (62.17%).
Cite
CITATION STYLE
Ghoul, D., & Lejeune, G. (2019). MICHAEL: Mining character-level patterns for arabic dialect identification (madar challenge). In ACL 2019 - 4th Arabic Natural Language Processing Workshop, WANLP 2019 - Proceedings of the Workshop (pp. 229–233). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w19-4627
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.