Fixing translation divergences in parallel corpora for neural MT

N/ACitations
Citations of this article
107Readers
Mendeley users who have this article in their library.

Abstract

Corpus-based approaches to machine translation rely on the availability of clean parallel corpora. Such resources are scarce, and because of the automatic processes involved in their preparation, they are often noisy. This paper describes an unsupervised method for detecting translation divergences in parallel sentences. We rely on a neural network that computes cross-lingual sentence similarity scores, which are then used to effectively filter out divergent translations. Furthermore, similarity scores predicted by the network are used to identify and fix some partial divergences, yielding additional parallel segments. We evaluate these methods for English-French and English-German machine translation tasks, and show that using filtered/corrected corpora actually improves MT performance.

Cite

CITATION STYLE

APA

Pham, M. Q., Crego, J., Senellart, J., & Yvon, F. (2018). Fixing translation divergences in parallel corpora for neural MT. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018 (pp. 2967–2973). Association for Computational Linguistics. https://doi.org/10.18653/v1/d18-1328

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free