Visual analysis of attention-based end-to-end speech recognition

  • Lim S
  • Goo J
  • Kim H
N/ACitations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

An end-to-end speech recognition model consisting of a single integrated neural network model was recently proposed. The end-to-end model does not need several training steps, and its structure is easy to understand. However, it is difficult to understand how the model recognizes speech internally. In this paper, we visualized and analyzed the attention-based end-to-end model to elucidate its internal mechanisms. We compared the acoustic model of the BLSTM-HMM hybrid model with the encoder of the end-to-end model, and visualized them using t-SNE to examine the difference between neural network layers. As a result, we were able to delineate the difference between the acoustic model and the end-to-end model encoder. Additionally, we analyzed the decoder of the end-to-end model from a language model perspective. Finally, we found that improving end-to-end model decoder is necessary to yield higher performance.

Cite

CITATION STYLE

APA

Lim, S., Goo, J., & Kim, H. (2019). Visual analysis of attention-based end-to-end speech recognition. Phonetics and Speech Sciences, 11(1), 41–49. https://doi.org/10.13064/ksss.2019.11.1.041

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free