On the Prediction Network Architecture in RNN-T for ASR

N/ACitations
Citations of this article
14Readers
Mendeley users who have this article in their library.

Abstract

RNN-T models have gained popularity in the literature and in commercial systems because of their competitiveness and capability of operating in online streaming mode. In this work, we conduct an extensive study comparing several prediction network architectures for both monotonic and original RNN-T models. We compare 4 types of prediction networks based on a common state-of-the-art Conformer encoder and report results obtained on Librispeech and an internal medical conversation data set. Our study covers both offline batch-mode and online streaming scenarios. In contrast to some previous works, our results show that Transformer does not always outperform LSTM when used as prediction network along with Conformer encoder. Inspired by our scoreboard, we propose a new simple prediction network architecture, N-Concat, that outperforms the others in our online streaming benchmark. Transformer and n-gram reduced architectures perform very similarly yet with some important distinct behaviour in terms of previous context. Overall we obtained up to 4.1 % relative WER improvement compared to our LSTM baseline, while reducing prediction network parameters by nearly an order of magnitude (8.4 times).

Cite

CITATION STYLE

APA

Albesano, D., Andrés-Ferrer, J., Ferri, N., & Zhan, P. (2022). On the Prediction Network Architecture in RNN-T for ASR. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH (Vol. 2022-September, pp. 2093–2097). International Speech Communication Association. https://doi.org/10.21437/Interspeech.2022-10954

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free