Abstract
Despite impressive performance and expanded language coverage in recent multilingual machine translation (MMT) systems, most African, and particularly East African, languages remain severely underrepresented in natural language processing (NLP) benchmarks, corpora, and state-of-the-art models (SOTA). With more than 2,000 languages spoken across Africa, current MT resources targets only a small fraction, leaving many languages under-served. To address this gap, we introduce AfriMMT-EA, the first large-scale multilingual MT dataset covering 53 East African languages across diverse domains. We expand low resource coverage by providing high-quality parallel data for many languages with no prior digital presence, including 23 new Kenyan and Tanzanian language pairs. We fine-tune Gemma-3-270M and Gemma-3-1B, to create our regionally adapted models Safari-270M and Safari-1B observing consistent translation quality improvements over strong off-the-shelf baselines, with the 1B model consistently outperforming the 270M variant across most languages. We release the dataset, trained models, and tools to lower barriers for researchers and communities, supporting more inclusive, robust, and culturally grounded MT research. All artifacts are publicly available at
Cite
CITATION STYLE
Etori, N. A., Ezema, K., Robinson, N. R., David, D., Kondoro, A. M., Makori, E. O., … Gini, M. L. (2026). AfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages. In 19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 (pp. 3459–3492). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2026.findings-eacl.179
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.