Estimating nonnegative matrix model activations with deep neural networks to increase perceptual speech quality

  • Williamson D
  • Wang Y
  • Wang D
21Citations
Citations of this article
36Readers
Mendeley users who have this article in their library.
Get full text

Abstract

As a means of speech separation, time-frequency masking applies a gain function to the time-frequency representation of noisy speech. On the other hand, nonnegative matrix factorization (NMF) addresses separation by linearly combining basis vectors from speech and noise models to approximate noisy speech. This paper presents an approach for improving the perceptual quality of speech separated from background noise at low signal-to-noise ratios. An ideal ratio mask is estimated, which separates speech from noise with reasonable sound quality. A deep neural network then approximates clean speech by estimating activation weights from the ratio-masked speech, where the weights linearly combine elements from a NMF speech model. Systematic comparisons using objective metrics, including the perceptual evaluation of speech quality, show that the proposed algorithm achieves higher speech quality than related masking and NMF methods. In addition, a listening test was performed and its results show that the output of the proposed algorithm is preferred over the comparison systems in terms of speech quality.

Cite

CITATION STYLE

APA

Williamson, D. S., Wang, Y., & Wang, D. (2015). Estimating nonnegative matrix model activations with deep neural networks to increase perceptual speech quality. The Journal of the Acoustical Society of America, 138(3), 1399–1407. https://doi.org/10.1121/1.4928612

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free