Interpreting Predictions of NLP Models

30Citations
Citations of this article
252Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Although neural NLP models are highly expressive and empirically successful, they also systematically fail in counterintuitive ways and are opaque in their decision-making process. This tutorial will provide a background on interpretation techniques, i.e., methods for explaining the predictions of NLP models. We will first situate example-specific interpretations in the context of other ways to understand models (e.g., probing, dataset analyses). Next, we will present a thorough study of example-specific interpretations, including saliency maps, input perturbations (e.g., LIME, input reduction), adversarial attacks, and influence functions. Alongside these descriptions, we will walk through source code that creates and visualizes interpretations for a diverse set of NLP tasks. Finally, we will discuss open problems in the field, e.g., evaluating, extending, and improving interpretation methods. The tutorial slides and the accompanying code is available online at https://www.ericswallace.com/interpretability.

Cite

CITATION STYLE

APA

Wallace, E., Gardner, M., & Singh, S. (2020). Interpreting Predictions of NLP Models. In EMNLP 2020 - Conference on Empirical Methods in Natural Language Processing, Tutorial Abstracts (pp. 20–23). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/P17

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free