Abstract
Part of Speech (POS) tagging is one of the most common techniques used in natural language processing (NLP) applications and corpus linguistics. Various POS tagging tools have been developed for Arabic. These taggers differ in several aspects, such as in their modeling techniques, tag sets and training and testing data. In this paper we conduct a comparative study of five Arabic POS taggers, namely: Stanford Arabic, CAMeL Tools, Farasa, MADAMIRA and Arabic Linguistic Pipeline (ALP) which examine their performance using text samples from Saudi novels. The testing data has been extracted from different novels that represent different types of narrations. The main result we have obtained indicates that the ALP tagger performs better than others in this particular case, and that Adjective is the most frequent mistagged POS type as compared to Noun and Verb.
Author supplied keywords
Cite
CITATION STYLE
Alluhaibi, R., Alfraidi, T., Abdeen, M. A. R., & Yatimi, A. (2021). A Comparative Study of Arabic Part of Speech Taggers Using Literary Text Samples from Saudi Novels. Information (Switzerland), 12(12). https://doi.org/10.3390/info12120523
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.