Abstract
We propose a low-memory-bandwidth, high-efficiency VLSI architecture for 60-k word real-time continuous speech recognition. Our architecture includes a cache architecture using the locality of speech recognition, beam pruning using a dynamic threshold, two-stage language model searching, a parallel Gaussian Mixture Model (GMM) architecture based on the mixture level and frame level, a parallel Viterbi architecture, and pipeline operation between Viterbi transition and GMM processing. Results show that our architecture achieves 88.24% required frequency reduction (66.74 MHz) and 84.04% memory bandwidth reduction (549.91MB/s) for real-time 60-k word continuous speech recognition. Copyright © 2011 The Institute of Electronics, Information and Communication Engineers.
Author supplied keywords
Cite
CITATION STYLE
Noguchi, H., Miura, K., Fujinaga, T., Sugahara, T., Kawaguchi, H., & Yoshimoto, M. (2011). VLSI architecture of GMM processing and viterbi decoder for 60,000-word real-time continuous speech recognition. IEICE Transactions on Electronics, E94-C(4), 458–467. https://doi.org/10.1587/transele.E94.C.458
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.