Massively Multilingual Instruction-Following Information Extraction

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The literature on information extraction (IE) has mostly centered around a selected few languages, hindering their applications on multilingual corpora. In this work, we introduce MASSIE - a comprehensive collection for instruction-following multilingual IE that standardizes and unifies 215 manually annotated datasets, covering 96 typologically diverse languages from 18 language families. Based on MASSIE, we conduct empirical studies on few-shot in-context learning and report important factors that either positively or negatively affect LLMs' performance in multilingual IE, covering 21 LLMs sizing from 0.5B to 72B. Additionally, we introduce LF1 - a structure-aware metric that captures partially matched spans, resolving the conservativeness of standard exact matching scheme which overpenalizes LLMs' predictions. Overall, our results signify that multilingual IE remains very challenging for existing LLMs, especially on complex tasks involving relations and events. In addition, performance gap is extremely large among high- and low-performing languages, but the group of similar-performing languages largely overlap between different LLMs, suggesting a shared performance bias in current LLMs.

Cite

CITATION STYLE

APA

Le, T., Nguyen, H. H., Tuan, L. A., & Nguyen, T. H. (2025). Massively Multilingual Instruction-Following Information Extraction. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 3542–3585). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.182

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free