BLONDE: An Automatic Evaluation Metric for Document-level Machine Translation

N/ACitations
Citations of this article
60Readers
Mendeley users who have this article in their library.

Abstract

Standard automatic metrics, e.g., BLEU, are not reliable for document-level MT evaluation. They can neither distinguish document-level improvements in translation quality from sentence-level ones, nor identify the discourse phenomena that cause context-agnostic translations. This paper introduces a novel automatic metric BLONDE to widen the scope of automatic MT evaluation from the sentence to the document level. BLONDE takes discourse coherence into consideration by categorizing discourse-related spans and calculating the similarity-based F1 measure of categorized spans. We conduct extensive comparisons on a newly constructed document-level translation dataset BWB. The experimental results show that BLONDE possesses better selectivity and interpretability at the document-level, and is more sensitive to document-level nuances. In a large-scale human study, BLONDE also achieves significantly higher Pearson's r correlation with human judgments compared to previous metrics.

Cite

CITATION STYLE

APA

Jiang, Y. E., Liu, T., Ma, S., Zhang, D., Yang, J., Huang, H., … Zhou, M. (2022). BLONDE: An Automatic Evaluation Metric for Document-level Machine Translation. In NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference (pp. 1550–1565). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.naacl-main.111

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free