Data augmentation for low-Resource neural machine translation

N/ACitations
Citations of this article
467Readers
Mendeley users who have this article in their library.

Abstract

The quality of a Neural Machine Translation system depends substantially on the availability of sizable parallel corpora. For low-resource language pairs this is not the case, resulting in poor translation quality. Inspired by work in computer vision, we propose a novel data augmentation approach that targets low-frequency words by generating new sentence pairs containing rare words in new, synthetically created contexts. Experimental results on simulated low-resource settings show that our method improves translation quality by up to 2.9 BLEU points over the baseline and up to 3.2 BLEU over back-translation.

Cite

CITATION STYLE

APA

Fadaee, M., Bisazza, A., & Monz, C. (2017). Data augmentation for low-Resource neural machine translation. In ACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) (Vol. 2, pp. 567–573). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/P17-2090

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free