Inducing word and part-of-speech with pitman-yor hidden semi-markov models

27Citations
Citations of this article
107Readers
Mendeley users who have this article in their library.

Abstract

We propose a nonparametric Bayesian model for joint unsupervised word segmentation and part-of-speech tagging from raw strings. Extending a previous model for word segmentation, our model is called a Pitman-Yor Hidden Semi-Markov Model (PYHSMM) and considered as a method to build a class n-gram language model directly from strings, while integrating character and word level information. Experimental results on standard datasets on Japanese, Chinese and Thai revealed it outperforms previous results to yield the state-of-The-Art accuracies. This model will also serve to analyze a structure of a language whose words are not identified a priori.

Cite

CITATION STYLE

APA

Uchiumi, K., Tsukahara, H., & Mochihashi, D. (2015). Inducing word and part-of-speech with pitman-yor hidden semi-markov models. In ACL-IJCNLP 2015 - 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, Proceedings of the Conference (Vol. 1, pp. 1774–1782). Association for Computational Linguistics (ACL). https://doi.org/10.3115/v1/p15-1171

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free