Mixture of Languages: Improved Multilingual Encoders Through Language Grouping

0Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We propose Mixture of Languages (MoL), a new strategy to pretrain largely multilingual encoders. Recent work in this field has relied on training transformer encoders on a large amount of multilingual data, with all parameters shared across all languages, without studying how to optimally balance language transfer and interference to achieve better performance. To address this, MoL proposes to group languages based on their similarity, and add parallel, sparsely activated layers that process each group independently. This architecture allows MoL to boost language transfer while minimizing interference, without increasing the active parameter count. We show that MoL largely outperforms a dense counterpart trained with the same configuration, as well as MoE models and public multilingual encoders such as XLM-R or mBERT on downstream tasks.

Cite

CITATION STYLE

APA

Janeiro, J. M., Alastruey, B., Massa, F., Elbayad, M., Piwowarski, B., Gallinari, P., & Barrault, L. (2025). Mixture of Languages: Improved Multilingual Encoders Through Language Grouping. In EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 29707–29722). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.emnlp-main.1509

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free