A Comparative Study of Arabic Part of Speech Taggers Using Literary Text Samples from Saudi Novels

10Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

Abstract

Part of Speech (POS) tagging is one of the most common techniques used in natural language processing (NLP) applications and corpus linguistics. Various POS tagging tools have been developed for Arabic. These taggers differ in several aspects, such as in their modeling techniques, tag sets and training and testing data. In this paper we conduct a comparative study of five Arabic POS taggers, namely: Stanford Arabic, CAMeL Tools, Farasa, MADAMIRA and Arabic Linguistic Pipeline (ALP) which examine their performance using text samples from Saudi novels. The testing data has been extracted from different novels that represent different types of narrations. The main result we have obtained indicates that the ALP tagger performs better than others in this particular case, and that Adjective is the most frequent mistagged POS type as compared to Noun and Verb.

Cite

CITATION STYLE

APA

Alluhaibi, R., Alfraidi, T., Abdeen, M. A. R., & Yatimi, A. (2021). A Comparative Study of Arabic Part of Speech Taggers Using Literary Text Samples from Saudi Novels. Information (Switzerland), 12(12). https://doi.org/10.3390/info12120523

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free