Improving Vietnamese Word Segmentation and POS Tagging using MEM with Various Kinds of Resources

  • Tran O
  • Le C
  • Ha T
N/ACitations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

Word segmentation and POS tagging are two important problems included in many NLP tasks. They, however, have not drawn much attention of Vietnamese researchers all over the world. In this paper, we focus on the integration of advantages from several resourses to improve the accuracy of Vietnamese word segmentation as well as POS tagging task. For word segmentation, we propose a solution in which we try to utilize multiple knowledge resources including dictionary-based model, N-gram model, and named entity recognition model and then integrate them into a Maximum Entropy model. The result of experiments on a public corpus has shown its effectiveness in comparison with the best current models. We got 95.30% F1 measure. For POS tagging, motivated from Chinese research and Vietnamese characteristics, we present a new kind of features based on the idea of word composition. We call it morpheme-based features. Our experiments based on two POS-tagged corpora showed that morpheme-based features always give promising results. In the best case, we got 89.64% precision on a Vietnamese POS-tagged corpus when using Maximum Entropy model.

Cite

CITATION STYLE

APA

Tran, O. T., Le, C. A., & Ha, T. Q. (2010). Improving Vietnamese Word Segmentation and POS Tagging using MEM with Various Kinds of Resources. Journal of Natural Language Processing, 17(3), 41–60. https://doi.org/10.5715/jnlp.17.3_41

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free