A call for more rigor in unsupervised cross-lingual learning

52Citations
Citations of this article
191Readers
Mendeley users who have this article in their library.

Abstract

We review motivations, definition, approaches, and methodology for unsupervised crosslingual learning and call for a more rigorous position in each of them. An existing rationale for such research is based on the lack of parallel data for many of the world's languages. However, we argue that a scenario without any parallel data and abundant monolingual data is unrealistic in practice. We also discuss different training signals that have been used in previous work, which depart from the pure unsupervised setting. We then describe common methodological issues in tuning and evaluation of unsupervised cross-lingual models and present best practices. Finally, we provide a unified outlook for different types of research in this area (i.e., cross-lingual word embeddings, deep multilingual pretraining, and unsupervised machine translation) and argue for comparable evaluation of these models.

Cite

CITATION STYLE

APA

Artetxe, M., Ruder, S., Yogatama, D., Labaka, G., & Agirre, E. (2020). A call for more rigor in unsupervised cross-lingual learning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 7375–7388). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.acl-main.658

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free