Abstract
We attempted to estimate subjective scores of the Japanese Diagnostic Rhyme Test (DRT), a two-to-one forced selection speech intel- ligibility test. We used automatic speech recognizers with language models that force one of thewords in theword-pair,mimicking the human recogni- tion process of theDRT. Initial testingwas done using speaker-independent models, and they showed significantly lower scores than subjective scores. The acoustic models were then adapted to each of the speakers in the corpus, and then adapted to noise at a specified SNR. Three different types of noise were tested: white noise, multi-talker (babble) noise, and pseudo- speech noise. The match between subjective and estimated scores improved significantly with noise-adapted models compared to speaker-independent models and the speaker-adapted models, when the adapted noise level and the tested level match. However, when SNR conditions do not match, the recognition scores degraded especially when tested SNR conditions were higher than the adapted noise level. Accordingly, we adapted the models to mixed levels of noise, i.e., multi-condition training. The adapted models now showed relatively high intelligibility matching subjective intelligibil- ity performance over all levels of noise. The correlation between subjec- tive and estimated intelligibility scores increased to 0.94 with multi-talker noise, 0.93 with white noise, and 0.89 with pseudo-speech noise, while the root mean square error (RMSE) reduced frommore than 40 to 13.10, 13.05 and 16.06, respectively.
Cite
CITATION STYLE
Kondo, K. (2011). Estimation of Speech Intelligibility Using Perceptual Speech Quality Scores. In Speech and Language Technologies. InTech. https://doi.org/10.5772/16822
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.