Abstract
The development of neural models has greatly improved the performance of machine translation, but these methods require large-scale parallel data, which can be difficult to obtain for low-resource language pairs. To address this issue, this research employs a pre-Trained multilingual model and fine-Tunes it by using a small bilingual dataset. Additionally, two data-Augmentation strategies are proposed to generate new training data: (i) back-Translation with the dataset from the source language? (ii) data augmentation via the English pivot language. The proposed approach is applied to the Khmer-Vietnamese machine translation. Experimental results show that our proposed approach outperforms the Google Translator model by 5.3% in terms of BLEU score on a test set of 2,000 Khmer-Vietnamese sentence pairs.
Author supplied keywords
Cite
CITATION STYLE
Quoc, T. N., Thanh, H. L., & Van, H. P. (2023). Khmer-Vietnamese Neural Machine Translation Improvement Using Data Augmentation Strategies. Informatica (Slovenia), 47(3), 349–360. https://doi.org/10.31449/inf.v47i3.4761
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.