Introspection for convolutional automatic speech recognition

Andreas Krug; Sebastian Stober

Conference ProceedingsOPEN ACCESS

Introspection for convolutional automatic speech recognition

EMNLP 2018 - 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Proceedings of the 1st Workshop (2018) 187-199

DOI: 10.18653/v1/w18-5421

12Citations

80Readers

Abstract

Artificial Neural Networks (ANNs) have experienced great success in the past few years. The increasing complexity of these models leads to less understanding about their decision processes. Therefore, introspection techniques have been proposed, mostly for images as input data. Patterns or relevant regions in images can be intuitively interpreted by a human observer. This is not the case for more complex data like speech recordings. In this work, we investigate the application of common introspection techniques from computer vision to an Automatic Speech Recognition (ASR) task. To this end, we use a model similar to image classification, which predicts letters from spectrograms. We show difficulties in applying image introspection to ASR. To tackle these problems, we propose normalized averaging of aligned inputs (NAvAI): a data-driven method to reveal learned patterns for prediction of specific classes. Our method integrates information from many data examples through local introspection techniques for Convolutional Neural Networks (CNNs). We demonstrate that our method provides better interpretability of letter-specific patterns than existing methods.

Cite

CITATION STYLE

APA

Krug, A., & Stober, S. (2018). Introspection for convolutional automatic speech recognition. In EMNLP 2018 - 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Proceedings of the 1st Workshop (pp. 187–199). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-5421

Introspection for convolutional automatic speech recognition

Abstract

Cite

Register to see more suggestions