Nonparametric Bayesian Semi-supervised Word Segmentation

  • Fujii R
  • Domoto R
  • Mochihashi D
N/ACitations
Citations of this article
89Readers
Mendeley users who have this article in their library.

Abstract

This paper presents a novel hybrid generative/discriminative model of word segmentation based on nonparametric Bayesian methods. Unlike ordinary discriminative word segmentation which relies only on labeled data, our semi-supervised model also leverages a huge amounts of unlabeled text to automatically learn new “words”, and further constrains them by using a labeled data to segment non-standard texts such as those found in social networking services.Specifically, our hybrid model combines a discriminative classifier (CRF; Lafferty et al. (2001) and unsupervised word segmentation (NPYLM; Mochihashi et al. (2009)), with a transparent exchange of information between these two model structures within the semi-supervised framework (JESS-CM; Suzuki and Isozaki (2008)). We confirmed that it can appropriately segment non-standard texts like those in Twitter and Weibo and has nearly state-of-the-art accuracy on standard datasets in Japanese, Chinese, and Thai.

Cite

CITATION STYLE

APA

Fujii, R., Domoto, R., & Mochihashi, D. (2017). Nonparametric Bayesian Semi-supervised Word Segmentation. Transactions of the Association for Computational Linguistics, 5, 179–189. https://doi.org/10.1162/tacl_a_00054

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free