Abstract
Translating Kannada-English (Kn-En) code-mixed text is a challenging task due to the limited availability of Kannada language resources and the inherent complexity of the dataset. This study evaluates the effectiveness of the sentence transformer model, utilizing pre-trained multilingual MPNet and Bidirectional Encoder Representations from Transformers (BERT) architectures, in generating sentence embeddings to enhance translation accuracy. It encodes both code-mixed sentences and their corresponding Kannada translations into high-dimensional embeddings. By employing cosine similarity, it maps input sentences to their closest translations, encoding 2000 code-mixed sentences and their translations using both the MPNet and BERT models. The findings indicate that the MPNet model proved to be more effective, achieving a model accuracy of 98%, compared to BERT's 88%. Moreover, MPNet outperformed BERT in terms of Bilingual Evaluation Understudy (BLEU) and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) scores, attaining 85.0 and 80.0, respectively, while BERT scored 65.3 and 58.7. These results highlight the advanced capabilities of MPNet in translating code-mixed languages and its potential applicability to a broader range of multilingual Natural Language Processing (NLP) tasks.
Author supplied keywords
Cite
CITATION STYLE
Rohith, H. P., Kumar, L., Kavitha, S., Karunakara, R. B., & Inchara, K. P. (2025). Performance Analysis of Effective Retrieval of Kannada Translations in Code-Mixed Sentences using BERT and MPnet. Engineering, Technology and Applied Science Research, 15(1), 19109–19114. https://doi.org/10.48084/etasr.9013
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.