Exploring morphology-aware tokenization: A case study on Spanish language modeling

3Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

This paper investigates to what extent the integration of morphological information can improve subword tokenization and thus also language modeling performance. We focus on Spanish, a language with fusional morphology, where subword segmentation can benefit from linguistic structure. Instead of relying on purely data-driven strategies like Byte Pair Encoding (BPE), we explore a linguistically grounded approach: training a tokenizer on morphologically segmented data. To do so, we develop a semi-supervised segmentation model for Spanish, building gold-standard datasets to guide and evaluate it. We then use this tokenizer to pre-train a masked language model and assess its performance on several downstream tasks. Our results show improvements over a baseline with a standard tokenizer, supporting our hypothesis that morphology-aware tokenization offers a viable and principled alternative for improving language modeling.

Cite

CITATION STYLE

APA

García, A. T., Przybyla, P., & Wanner, L. (2025). Exploring morphology-aware tokenization: A case study on Spanish language modeling. In EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 30505–30518). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.emnlp-main.1552

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free