Impact of phase estimation on single-channel speech separation based on time-frequency masking

  • Mayer F
  • Williamson D
  • Mowlaee P
  • et al.
27Citations
Citations of this article
33Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Time-frequency masking is a common solution for the single-channel source separation (SCSS) problem where the goal is to find a time-frequency mask that separates the underlying sources from an observed mixture. An estimated mask is then applied to the mixed signal to extract the desired signal. During signal reconstruction, the time-frequency–masked spectral amplitude is combined with the mixture phase. This article considers the impact of replacing the mixture spectral phase with an estimated clean spectral phase combined with the estimated magnitude spectrum using a conventional model-based approach. As the proposed phase estimator requires estimated fundamental frequency of the underlying signal from the mixture, a robust pitch estimator is proposed. The upper-bound clean phase results show the potential of phase-aware processing in single-channel source separation. Also, the experiments demonstrate that replacing the mixture phase with the estimated clean spectral phase consistently improves perceptual speech quality, predicted speech intelligibility, and source separation performance across all signal-to-noise ratio and noise scenarios.

Cite

CITATION STYLE

APA

Mayer, F., Williamson, D. S., Mowlaee, P., & Wang, D. (2017). Impact of phase estimation on single-channel speech separation based on time-frequency masking. The Journal of the Acoustical Society of America, 141(6), 4668–4679. https://doi.org/10.1121/1.4986647

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free