Morphological disambiguation from stemming data

Antoine Nzeyimana

Conference ProceedingsOPEN ACCESS

Morphological disambiguation from stemming data

Nzeyimana A

COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference (2020) 4649-4660

DOI: 10.18653/v1/2020.coling-main.409

5Citations

71Readers

Abstract

Morphological analysis and disambiguation is an important task and a crucial preprocessing step in natural language processing of morphologically rich languages. Kinyarwanda, a morphologically rich language, currently lacks tools for automated morphological analysis. While linguistically curated finite state tools can be easily developed for morphological analysis, the morphological richness of the language allows many ambiguous analyses to be produced, requiring effective disambiguation. In this paper, we propose learning to morphologically disambiguate Kinyarwanda verbal forms from a new stemming dataset collected through crowd-sourcing. Using feature engineering and a feed-forward neural network based classifier, we achieve about 89% non-contextualized disambiguation accuracy. Our experiments reveal that inflectional properties of stems and morpheme association rules are the most discriminative features for disambiguation.

Cite

CITATION STYLE

APA

Nzeyimana, A. (2020). Morphological disambiguation from stemming data. In COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference (pp. 4649–4660). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.coling-main.409

Morphological disambiguation from stemming data

Abstract

Cite

Register to see more suggestions