Abstract
This paper describes the system submitted by IITP-MT team to Computational Approaches to Linguistic Code-Switching (CALCS 2021) shared task on MT for English → Hinglish. We submit a neural machine translation (NMT) system which is trained on the synthetic code-mixed (cm) English-Hinglish parallel corpus. We propose an approach to create code-mixed parallel corpus from a clean parallel corpus in an unsupervised manner. It is an alignment based approach and we do not use any linguistic resources for explicitly marking any token for code-switching. We also train NMT model on the gold corpus provided by the workshop organizers augmented with the generated synthetic code-mixed parallel corpus. The model trained over the generated synthetic cm data achieves 10.09 BLEU points over the given test set.
Cite
CITATION STYLE
Appicharla, R., Gupta, K. K., Ekbal, A., & Bhattacharyya, P. (2021). IITP-MT at CALCS2021: English to Hinglish Neural Machine Translation using Unsupervised Synthetic Code-Mixed Parallel Corpus. In Computational Approaches to Linguistic Code-Switching, CALCS 2021 - Proceedings of the 5th Workshop (pp. 31–35). Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-056-4_005
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.