VLSI architecture of GMM processing and viterbi decoder for 60,000-word real-time continuous speech recognition

6Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

We propose a low-memory-bandwidth, high-efficiency VLSI architecture for 60-k word real-time continuous speech recognition. Our architecture includes a cache architecture using the locality of speech recognition, beam pruning using a dynamic threshold, two-stage language model searching, a parallel Gaussian Mixture Model (GMM) architecture based on the mixture level and frame level, a parallel Viterbi architecture, and pipeline operation between Viterbi transition and GMM processing. Results show that our architecture achieves 88.24% required frequency reduction (66.74 MHz) and 84.04% memory bandwidth reduction (549.91MB/s) for real-time 60-k word continuous speech recognition. Copyright © 2011 The Institute of Electronics, Information and Communication Engineers.

Cite

CITATION STYLE

APA

Noguchi, H., Miura, K., Fujinaga, T., Sugahara, T., Kawaguchi, H., & Yoshimoto, M. (2011). VLSI architecture of GMM processing and viterbi decoder for 60,000-word real-time continuous speech recognition. IEICE Transactions on Electronics, E94-C(4), 458–467. https://doi.org/10.1587/transele.E94.C.458

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free