Reference-Free Word- and Sentence-Level Translation Evaluation with Token-Matching Metrics

Christoph Wolfgang Leiter

Conference ProceedingsOPEN ACCESS

Reference-Free Word- and Sentence-Level Translation Evaluation with Token-Matching Metrics

Leiter C

Eval4NLP 2021 - Evaluation and Comparison of NLP Systems, Proceedings of the 2nd Workshop (2021) 157-164

DOI: 10.26615/978-954-452-056-4_016

6Citations

33Readers

Get full text

Abstract

Many modern machine translation evaluation metrics like BERTScore, BLEURT, COMET, MonoTransquest or XMoverScore are based on black-box language models. Hence, it is difficult to explain why these metrics return certain scores. This year’s Eval4NLP shared task tackles this challenge by searching for methods that can extract feature importance scores that correlate well with human word-level error annotations. In this paper we show that unsupervised metrics that are based on token-matching can intrinsically provide such scores. The submitted system interprets the similarities of the contextualized word-embeddings that are used to compute (X)BERTScore as word-level importance scores. We make our code available.

Cite

CITATION STYLE

APA

Leiter, C. W. (2021). Reference-Free Word- and Sentence-Level Translation Evaluation with Token-Matching Metrics. In Eval4NLP 2021 - Evaluation and Comparison of NLP Systems, Proceedings of the 2nd Workshop (pp. 157–164). Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-056-4_016

Reference-Free Word- and Sentence-Level Translation Evaluation with Token-Matching Metrics

Abstract

Cite

Register to see more suggestions