Abstract
We propose a new method for source separation by synthesizing the source from a speech mixture corrupted by various environmental noise. Unlike traditional source separation methods which estimate the source from the mixture as a replica of the original source (e.g. by solving an inverse problem), our proposed method is a synthesis-based approach which aims to generate a new signal (i.e. 'fake' source) that sounds similar to the original source. The proposed system has an encoder-decoder topology, where the encoder predicts intermediate-level features from the mixture, i.e. Mel-spectrum of the target source, using a hybrid recurrent and hourglass network, while the decoder is a state-of-the-art WaveNet speech synthesis network conditioned on the Mel-spectrum, which directly generates time-domain samples of the sources. Both objective and subjective evaluations were performed on the synthesized sources, and show great advantages of our proposed method for high-quality speech source separation and generation.
Author supplied keywords
Cite
CITATION STYLE
Liu, Q., Jackson, P. J. B., & Wang, W. (2019). A Speech Synthesis Approach for High Quality Speech Separation and Generation. IEEE Signal Processing Letters, 26(12), 1872–1876. https://doi.org/10.1109/LSP.2019.2951894
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.