Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions

34Citations
Citations of this article
53Readers
Mendeley users who have this article in their library.
Get full text

Abstract

This paper focuses on single-channel semi-supervised speech enhancement. We learn a speaker-independent deep generative speech model using the framework of variational autoencoders. The noise model remains unsupervised because we do not assume prior knowledge of the noisy recording environment. In this context, our contribution is to propose a noise model based on alpha-stable distributions, instead of the more conventional Gaussian non-negative matrix factorization approach found in previous studies. We develop a Monte Carlo expectation-maximization algorithm for estimating the model parameters at test time. Experimental results show the superiority of the proposed approach both in terms of perceptual quality and intelligibility of the enhanced speech signal.

Cite

CITATION STYLE

APA

Leglaive, S., Simsekli, U., Liutkus, A., Girin, L., & Horaud, R. (2019). Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings (Vol. 2019-May, pp. 541–545). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/ICASSP.2019.8682546

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free