Abstract
This paper describes the submissions of the “Marian” team to the WNMT 2018 shared task. We investigate combinations of teacher-student training, low-precision matrix products, auto-tuning and other methods to optimize the Transformer model on GPU and CPU. By further integrating these methods with the new averaging attention networks, a recently introduced faster Transformer variant, we create a number of high-quality, high-performance models on the GPU and CPU, dominating the Pareto frontier for this shared task.
Cite
CITATION STYLE
Junczys-Dowmunt, M., Heafield, K., Hoang, H., Grundkiewicz, R., & Aue, A. (2018). Marian: Cost-effective High-Quality Neural Machine Translation in C++. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 129–135). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-2716
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.