Abstract
We present a very simple method for parallel text cleaning of low-resource languages, based on projection of word embeddings trained on large monolingual corpora in high-resource languages. In spite of its simplicity, we approach the strong baseline system in the downstream machine translation evaluation.
Cite
CITATION STYLE
APA
Kurfalı, M., & Östling, R. (2019). Noisy parallel corpus filtering through projected word embeddings. In WMT 2019 - 4th Conference on Machine Translation, Proceedings of the Conference (Vol. 3, pp. 277–281). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w19-5438
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.
Already have an account? Sign in
Sign up for free