Asking without telling: Exploring latent ontologies in contextual representations

Julian Michael; Jan A. Botha; Ian Tenney

Conference ProceedingsOPEN ACCESS

Asking without telling: Exploring latent ontologies in contextual representations

EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (2020) 6792-6812

DOI: 10.18653/v1/2020.emnlp-main.552

28Citations

100Readers

Abstract

The success of pretrained contextual encoders, such as ELMo and BERT, has brought a great deal of interest in what these models learn: do they, without explicit supervision, learn to encode meaningful notions of linguistic structure? If so, how is this structure encoded? To investigate this, we introduce latent subclass learning (LSL): a modification to classifier-based probing that induces a latent categorization (or ontology) of the probe's inputs. Without access to fine-grained gold labels, LSL extracts emergent structure from input representations in an interpretable and quantifiable form. In experiments, we find strong evidence of familiar categories, such as a notion of personhood in ELMo, as well as novel ontological distinctions, such as a preference for fine-grained semantic roles on core arguments. Our results provide unique new evidence of emergent structure in pretrained encoders, including departures from existing annotations which are inaccessible to earlier methods.

Cite

CITATION STYLE

APA

Michael, J., Botha, J. A., & Tenney, I. (2020). Asking without telling: Exploring latent ontologies in contextual representations. In EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 6792–6812). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.emnlp-main.552

Asking without telling: Exploring latent ontologies in contextual representations

Abstract

Cite

Register to see more suggestions