High-Fidelity and Pitch-Controllable Neural Vocoder Based on Unified Source-Filter Networks

8Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

We introduce unified source-filter generative adversarial networks (uSFGAN), a waveform generative model conditioned on acoustic features, which represents the source-filter architecture in a generator network. Unlike the previous neural-based source-filter models in which parametric signal process modules are combined with neural networks, our approach enables unified optimization of both the source excitation generation and resonance filtering parts to achieve higher sound quality. In the uSFGAN framework, several specific regularization losses are proposed to enable the source excitation generation part to output reasonable source excitation signals. Both objective and subjective experiments are conducted, and the results demonstrate that the proposed uSFGAN achieves comparable sound quality to HiFi-GAN in the speech reconstruction task and outperforms WORLD in the $\text{F}_{0}$ transformation task. Moreover, we argue that the $\text{F}_{0}$-driven mechanism and the inductive bias obtained by source-filter modeling improve the robustness against unseen $\text{F}_{0}$ in training as shown by the results of experimental evaluations. Audio samples are available at our demo site at https://chomeyama.github.io/PitchControllableNeuralVocoder-Demo/.

Cite

CITATION STYLE

APA

Yoneyama, R., Wu, Y. C., & Toda, T. (2023). High-Fidelity and Pitch-Controllable Neural Vocoder Based on Unified Source-Filter Networks. IEEE/ACM Transactions on Audio Speech and Language Processing, 31, 3717–3729. https://doi.org/10.1109/TASLP.2023.3313410

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free