Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs

1Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-English data and utilizing a custom multilingual tokenizer, our models address the limitations of existing Large Language Models (LLMs) that predominantly focus on English or a few high-resource languages. We detail the models' development principles, i.e., data composition, tokenizer optimization, and training methodologies. The models demonstrate strong performance across multilingual benchmarks, as evidenced by their performance on European versions of ARC, HellaSwag, and TruthfulQA.

Cite

CITATION STYLE

APA

Ali, M., Fromm, M., Thellmann, K., Ebert, J., Weber, A. A., Rutmann, R., … Flores-Herr, N. (2025). Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs. In Frontiers in Artificial Intelligence and Applications (Vol. 413, pp. 4321–4329). IOS Press BV. https://doi.org/10.3233/FAIA251328

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free