Abstract
We introduce unified source-filter generative adversarial networks (uSFGAN), a waveform generative model conditioned on acoustic features, which represents the source-filter architecture in a generator network. Unlike the previous neural-based source-filter models in which parametric signal process modules are combined with neural networks, our approach enables unified optimization of both the source excitation generation and resonance filtering parts to achieve higher sound quality. In the uSFGAN framework, several specific regularization losses are proposed to enable the source excitation generation part to output reasonable source excitation signals. Both objective and subjective experiments are conducted, and the results demonstrate that the proposed uSFGAN achieves comparable sound quality to HiFi-GAN in the speech reconstruction task and outperforms WORLD in the $\text{F}_{0}$ transformation task. Moreover, we argue that the $\text{F}_{0}$-driven mechanism and the inductive bias obtained by source-filter modeling improve the robustness against unseen $\text{F}_{0}$ in training as shown by the results of experimental evaluations. Audio samples are available at our demo site at https://chomeyama.github.io/PitchControllableNeuralVocoder-Demo/.
Author supplied keywords
Cite
CITATION STYLE
Yoneyama, R., Wu, Y. C., & Toda, T. (2023). High-Fidelity and Pitch-Controllable Neural Vocoder Based on Unified Source-Filter Networks. IEEE/ACM Transactions on Audio Speech and Language Processing, 31, 3717–3729. https://doi.org/10.1109/TASLP.2023.3313410
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.