Abstract
Machine translation has advanced considerably in recent years yet producing accurate and natural translations remains challenging for languages with limited digital resources. The scarcity of high-quality parallel data particularly hampers the performance of models in low-resource languages. This study introduces a preference-based learning approach grounded in reinforcement learning principles to improve translation quality by incorporating human feedback during the fine-tuning of the multilingual mT5-Large model. Two optimization techniques are employed: Direct Preference Optimization, which aligns model outputs with human preferences through pairwise comparison, and Hypergeometric Gamma Reward, a novel method proposed in this work that generates reward signals based on semantic similarity scores computed using Sentence-BERT. These rewards are integrated into the translation model through reward-weighted loss optimization. The approach is evaluated on translations from Dogri, Kashmiri, and Konkani into English using two standard test sets, IN22-Gen and IN22-Conv. The enhanced model achieves BLEU scores of 49.81 and 57.21 for Dogri to English, 43.57 and 46.57 for Kashmiri to English, and 38.75 and 34.56 for Konkani to English on IN22-Gen and IN22-Conv, respectively. All reported improvements were validated through the Approximate Randomization Test and repeated-measures Analysis of Variance, both confirming that the observed gains are statistically significant. The statistical analysis further demonstrates that the combined Direct Preference Optimization and Hypergeometric Gamma Reward model consistently achieves the strongest improvements in both syntactic accuracy and semantic similarity across all datasets. Despite using significantly less training data than the state-of-the-art IndicTrans2 model, the proposed method delivers superior or comparable performance to transformer-based architectures and other large language models. These results demonstrate the effectiveness of integrating human feedback through the combined use of reinforcement learning–inspired preference learning and reward-based optimization, highlighting their potential to enhance translation quality for low-resource and underrepresented languages.
Author supplied keywords
Cite
CITATION STYLE
Rajagopalan Nair, A., Gupta, D., Paul, B., & Siva Bhavani, J. (2026). Enhancing Low-Resource Indian Language Machine Translation Using Large Language Models With Preference Optimization and Hypergeometric-Gamma Reward. IEEE Access, 14, 1641–1665. https://doi.org/10.1109/ACCESS.2025.3648480
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.