Low-Resource Grammatical Error Correction: Selective Data Augmentation with Round-Trip Machine Translation

N/ACitations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Supervised state-of-the-art methods for grammatical error correction require large amounts of parallel data for training. Due to lack of gold-labeled data, techniques that create synthetic training data have become popular. We show that models trained on synthetic data tend to correct a limited range of grammar and spelling mistakes that involve character-level changes, but perform poorly on (more complex) phenomena that require word-level changes. We propose to address the performance gap on such errors by generating synthetic data through selective data augmentation via round-trip machine translation. We show that the proposed technique, SeLex-RT, is capable of generating mistakes that are similar to those observed with language learners. Using the approach with two types of state-of-the-art learning frameworks and two low-resource languages (Russian and Ukrainian), we achieve substantial improvements, compared to training on synthetic data produced with standard techniques. Analysis of the output reveals that models trained on data noisified with the SeLex-RT approach are capable of making word-level changes and correct lexical errors common with language learners.

Cite

CITATION STYLE

APA

Gomez, F. P., & Rozovskaya, A. (2025). Low-Resource Grammatical Error Correction: Selective Data Augmentation with Round-Trip Machine Translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 25749–25770). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.1322

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free