An efficient algorithm for unsupervised word segmentation with branching entropy and MDL

14Citations
Citations of this article
2Readers
Mendeley users who have this article in their library.

Abstract

This paper proposes a fast and simple unsuper-vised word segmentation algorithm that utilizes the local predictability of adjacent character sequences, while searching for a least-effort representation of the data. The model uses branching entropy as a means of constraining the hypothesis space, in order to efficiently obtain a solution that minimizes the length of a two-part MDL code. An evaluation with corpora in Japanese, Thai, English, and the "CHILDES" corpus for research in language development reveals that the algorithm achieves an accuracy, comparable to that of the state-of-the-art methods in unsupervised word segmentation, in a significantly reduced computational time. © 2010 Association for Computational Linguistics.

Cite

CITATION STYLE

APA

Zhikov, V., Takamura, H., & Okumura, M. (2010). An efficient algorithm for unsupervised word segmentation with branching entropy and MDL. In EMNLP 2010 - Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 832–842).

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free