Abstract
We compare the performance of the APT and AutoPRF metrics for pronoun translation against a manually annotated dataset comprising human judgements as to the correctness of translations of the PROTEST test suite. Although there is some correlation with the human judgements, a range of issues limit the performance of the automated metrics. Instead, we recommend the use of semiautomatic metrics and test suites in place of fully automatic metrics.
Cite
CITATION STYLE
Guillou, L., & Hardmeier, C. (2018). Automatic reference-based evaluation of pronoun translation misses the point. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018 (pp. 4797–4802). Association for Computational Linguistics. https://doi.org/10.18653/v1/d18-1513
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.