Incorporating Residual and Normalization Layers into Analysis of Masked Language Models

N/ACitations
Citations of this article
88Readers
Mendeley users who have this article in their library.

Abstract

Transformer architecture has become ubiquitous in the natural language processing field. To interpret the Transformer-based models, their attention patterns have been extensively analyzed. However, the Transformer architecture is not only composed of the multi-head attention; other components can also contribute to Transformers' progressive performance. In this study, we extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization. Our analysis of Transformer-based masked language models shows that the token-to-token interaction performed via attention has less impact on the intermediate representations than previously assumed. These results provide new intuitive explanations of existing reports; for example, discarding the learned attention patterns tends not to adversely affect the performance. The codes of our experiments are publicly available.

Cite

CITATION STYLE

APA

Kobayashi, G., Kuribayashi, T., Yokoi, S., & Inui, K. (2021). Incorporating Residual and Normalization Layers into Analysis of Masked Language Models. In EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing, Proceedings (pp. 4547–4568). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.emnlp-main.373

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free