Rethinking the Value of Transformer Components

35Citations
Citations of this article
116Readers
Mendeley users who have this article in their library.

Abstract

Transformer becomes the state-of-the-art translation model, while it is not well studied how each intermediate component contributes to the model performance, which poses significant challenges for designing optimal architectures. In this work, we bridge this gap by evaluating the impact of individual component (sub-layer) in trained Transformer models from different perspectives. Experimental results across language pairs, training strategies, and model capacities show that certain components are consistently more important than the others. We also report a number of interesting findings that might help humans better analyze, understand and improve Transformer models. Based on these observations, we further propose a new training strategy that can improves translation performance by distinguishing the unimportant components in training.

Cite

CITATION STYLE

APA

Wang, W., & Tu, Z. (2020). Rethinking the Value of Transformer Components. In COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference (pp. 6019–6029). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.coling-main.529

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free