Abstract
Unsupervised learning of units (phonemes, words, phrases, etc.) is important to the design of statistical speech and NLP systems. This paper presents a general source-coding framework for inducing words from natural language text without word boundaries. An efficient search algorithm is developed to optimize the minimum description length (MDL) induction criterion. Despite some seemingly oversimplified modeling assumption, we achieved good results on several word induction problems.
Cite
CITATION STYLE
Yu, H. (2000). Unsupervised word induction using MDL criterion. In Proceedings of ISCSL. Retrieved from https://pdfs.semanticscholar.org/1d48/b919f8419f6f5a0abfcadf12d0caa7155b85.pdf
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.