Indonesian part of speech tagging using maximum entropy markov model on Indonesian manually tagged corpus

Denis Eka Cahyani; Winda Mustikaningtyas

Journal ArticleOPEN ACCESS

Indonesian part of speech tagging using maximum entropy markov model on Indonesian manually tagged corpus

IAES International Journal of Artificial Intelligence (2022) 11(1) 336-344

DOI: 10.11591/ijai.v11.i1.pp336-344

2Citations

14Readers

Abstract

This research discusses the development of a part of speech (POS) tagging system to solve the problem of word ambiguity. This paper presents a new method, namely maximum entropy markov model (MEMM) to solve word ambiguity on the Indonesian dataset. A manually labeled “Indonesian manually tagged corpus” was used as data. Furthermore, the corpus is processed using the entropy formula to obtain the weight of the value of the word being searched for, then calculating it into the MEMM Bigram and MEMM Trigram algorithms with the previously obtained rules to determine the part of speech (POS) tag that has the highest probability. The results obtained show POS tagging using the MEMM method has advantages over the methods used previously which used the same data. This paper improves a performance evaluation of research previously. The resulting average accuracy is 83.04% for the MEMM Bigram algorithm and 86.66% for the MEMM Trigram. The MEMM Trigram algorithm is better than the MEMM Bigram algorithm.

Author supplied keywords

Cite

CITATION STYLE

APA

Cahyani, D. E., & Mustikaningtyas, W. (2022). Indonesian part of speech tagging using maximum entropy markov model on Indonesian manually tagged corpus. IAES International Journal of Artificial Intelligence, 11(1), 336–344. https://doi.org/10.11591/ijai.v11.i1.pp336-344

Indonesian part of speech tagging using maximum entropy markov model on Indonesian manually tagged corpus

Abstract

Author supplied keywords

Cite

Register to see more suggestions