Abstract
Language tagging, a method whereby source and target inputs are prefixed with a unique language token, has become the de facto standard for conditioning Multilingual Neural Machine Translation (MNMT) models on specific language directions. This conditioning can manifest effective zero-shot translation abilities in MT models at scale for many languages. Expanding on previous work, we propose a novel method of language tagging for MNMT, injection, in which the embedded representation of a language token is concatenated to the input of every linear layer. We explore a variety of different tagging methods, with and without injection, showing that injection improves zero-shot translation performance with up to a 2+ BLEU score point gain for certain language directions in our dataset.
Cite
CITATION STYLE
Orten, J., Shurtz, A., Fulda, N., & Richardson, S. D. (2025). Neuron-Level Language Tag Injection Improves Zero-Shot Translation Performance. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 4, pp. 203–212). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-srw.13
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.