LLMs for Extremely Low-Resource Finno-Ugric Languages

3Citations
Citations of this article
13Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The advancement of large language models (LLMs) has predominantly focused on high-resource languages, leaving low-resource languages, such as those in the Finno-Ugric family, significantly underrepresented. This paper addresses this gap by focusing on Võro, Livonian, and Komi. We cover almost the entire cycle of LLM creation, from data collection to instruction tuning and evaluation. Our contributions include developing multilingual base and instruction-tuned models; creating evaluation benchmarks, including the SMUGRI-MT-BENCH multi-turn conversational benchmark; and conducting human evaluation. We intend for this work to promote linguistic diversity, ensuring that lesser-resourced languages can benefit from advancements in NLP.

Cite

CITATION STYLE

APA

Purason, T., Kuulmets, H. A., & Fishel, M. (2025). LLMs for Extremely Low-Resource Finno-Ugric Languages. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 6692–6712). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.373

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free