Towards better language representation in Natural Language Processing A multilingual dataset for text-level Grammatical Error Correction

1Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous G E C datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as G E C itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts — typically learner essays — rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual G E C studies.

Cite

CITATION STYLE

APA

Masciolini, A., Caines, A., De Clercq, O., Kruijsbergen, J., Kurfali, M., Muñoz Sánchez, R., … Zesch, T. (2025, May 15). Towards better language representation in Natural Language Processing A multilingual dataset for text-level Grammatical Error Correction. International Journal of Learner Corpus Research. John Benjamins Publishing Company. https://doi.org/10.1075/ijlcr.24033.mas

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free