From BERT's Point of View: Revealing the Prevailing Contextual Differences

N/ACitations
Citations of this article
42Readers
Mendeley users who have this article in their library.

Abstract

Though successfully applied in research and industry large pretrained language models of the BERT family are not yet fully understood. While much research in the field of BERTology has tested whether specific knowledge can be extracted from layer activations, we invert the popular probing design to analyze the prevailing differences and clusters in BERT's high dimensional space. By extracting coarse features from masked token representations and predicting them by probing models with access to only partial information we can apprehend the variation from 'BERT's point of view'. By applying our new methodology to different datasets we show how much the differences can be described by syntax but further how they are to a great extent shaped by the most simple positional information.

Cite

CITATION STYLE

APA

Schuster, C. M., & Hegelich, S. (2022). From BERT’s Point of View: Revealing the Prevailing Contextual Differences. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 1120–1138). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.findings-acl.89

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free