FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation

20Citations
Citations of this article
37Readers
Mendeley users who have this article in their library.

Abstract

We present FRMT, a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation. The dataset consists of professional translations from English into two regional variants each of Portuguese and Mandarin Chinese. Source documents are selected to enable detailed analysis of phenomena of interest, including lexically distinct terms and distractor terms. We explore automatic evaluation metrics for FRMT and validate their correlation with expert human evaluation across both region-matched and mismatched rating scenarios. Finally, we present a number of baseline models for this task, and offer guidelines for how researchers can train, evaluate, and compare their own models. Our dataset and evaluation code are publicly available: https://bit.ly/frmt-task.

Cite

CITATION STYLE

APA

Riley, P., Dozat, T., Botha, J. A., Garcia, X., Garrette, D., Riesa, J., … Constant, N. (2023). FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation. Transactions of the Association for Computational Linguistics, 11, 671–685. https://doi.org/10.1162/tacl_a_00568

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free