Khmer-Vietnamese Neural Machine Translation Improvement Using Data Augmentation Strategies

N/ACitations
Citations of this article
18Readers
Mendeley users who have this article in their library.

Abstract

The development of neural models has greatly improved the performance of machine translation, but these methods require large-scale parallel data, which can be difficult to obtain for low-resource language pairs. To address this issue, this research employs a pre-Trained multilingual model and fine-Tunes it by using a small bilingual dataset. Additionally, two data-Augmentation strategies are proposed to generate new training data: (i) back-Translation with the dataset from the source language? (ii) data augmentation via the English pivot language. The proposed approach is applied to the Khmer-Vietnamese machine translation. Experimental results show that our proposed approach outperforms the Google Translator model by 5.3% in terms of BLEU score on a test set of 2,000 Khmer-Vietnamese sentence pairs.

Cite

CITATION STYLE

APA

Quoc, T. N., Thanh, H. L., & Van, H. P. (2023). Khmer-Vietnamese Neural Machine Translation Improvement Using Data Augmentation Strategies. Informatica (Slovenia), 47(3), 349–360. https://doi.org/10.31449/inf.v47i3.4761

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free