QE4PE: Word-level Quality Estimation for Human Post-Editing

0Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Word-level quality estimation (QE) methods aim to detect erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the speed, quality, and editing choices of human post-editing remain understudied. In this study, we investigate the impact of word-level QE on machine translation (MT) post-editing in a realistic setting involving 42 professional post-editors across two translation directions. We compare four error-span highlight modalities, including supervised and uncertainty-based word-level QE methods, for identifying potential errors in the outputs of a state-of-the-art neural MT model. Post-editing effort and productivity are estimated from behavioral logs, while quality improvements are assessed by word- and segment-level human annotation. We find that domain, language and editors’ speed are critical factors in determining highlights’ effectiveness, with modest differences between human-made and automated QE highlights underlining a gap between accuracy and usability in professional workflows.

Cite

CITATION STYLE

APA

Sarti, G., Zouhar, V., Chrupała, G., Guerberof-Arenas, A., Nissim, M., & Bisazza, A. (2025). QE4PE: Word-level Quality Estimation for Human Post-Editing. Transactions of the Association for Computational Linguistics, 13, 1410–1435. https://doi.org/10.1162/TACL.a.46

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free