Statistical models for unsupervised, semi-supervised, and supervised transliteration mining

12Citations
Citations of this article
97Readers
Mendeley users who have this article in their library.

Abstract

We present a generative model that efficiently mines transliteration pairs in a consistent fashion in three different settings: unsupervised, semi-supervised, and supervised transliteration mining. The model interpolates two sub-models, one for the generation of transliteration pairs and one for the generation of non-transliteration pairs (i.e., noise). The model is trained on noisy unlabeled data using the EM algorithm. During training the transliteration submodel learns to generate transliteration pairs and the fixed non-transliteration model generates the noise pairs. After training, the unlabeled data is disambiguated based on the posterior probabilities of the two sub-models. We evaluate our transliteration mining system on data from a transliteration mining shared task and on parallel corpora. For three out of four language pairs, our system outperforms all semi-supervised and supervised systems that participated in the NEWS 2010 shared task. On word pairs extracted from parallel corpora with fewer than 2% transliteration pairs, our system achieves up to 86.7% F-measure with 77.9% precision and 97.8% recall.

Cite

CITATION STYLE

APA

Sajjad, H., Schmid, H., Fraser, A., & Schütze, H. (2017). Statistical models for unsupervised, semi-supervised, and supervised transliteration mining. Computational Linguistics, 43(2), 350–375. https://doi.org/10.1162/COLI_a_00286

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free