Abstract
The quality of a Neural Machine Translation system depends substantially on the availability of sizable parallel corpora. For low-resource language pairs this is not the case, resulting in poor translation quality. Inspired by work in computer vision, we propose a novel data augmentation approach that targets low-frequency words by generating new sentence pairs containing rare words in new, synthetically created contexts. Experimental results on simulated low-resource settings show that our method improves translation quality by up to 2.9 BLEU points over the baseline and up to 3.2 BLEU over back-translation.
Cite
CITATION STYLE
Fadaee, M., Bisazza, A., & Monz, C. (2017). Data augmentation for low-Resource neural machine translation. In ACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) (Vol. 2, pp. 567–573). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/P17-2090
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.