Cross-Lingual Voice Conversion with Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network

Tuan Vu Ho; Masato Akagi

Journal ArticleOPEN ACCESS

Cross-Lingual Voice Conversion with Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network

IEEE Access (2021) 9 47503-47515

DOI: 10.1109/ACCESS.2021.3063519

7Citations

12Readers

Abstract

This paper proposes a non-parallel cross-lingual voice conversion (CLVC) model that can mimic voice while continuously controlling speaker individuality on the basis of the variational autoencoder (VAE) and star generative adversarial network (StarGAN). Most studies on CLVC only focused on mimicking a particular speaker voice without being able to arbitrarily modify the speaker individuality. In practice, the ability to generate speaker individuality may be more useful than just mimicking voice. Therefore, the proposed model reliably extracts the speaker embedding from different languages using a VAE. An F0 injection method is also introduced into our model to enhance the F0 modeling in the cross-lingual setting. To avoid the over-smoothing degradation problem of the conventional VAE, the adversarial training scheme of the StarGAN is adopted to improve the training-objective function of the VAE in a CLVC task. Objective and subjective measurements confirm the effectiveness of the proposed model and F0 injection method. Furthermore, speaker-similarity measurement on fictitious voices reveal a strong linear relationship between speaker individuality and interpolated speaker embedding, which indicates that speaker individuality can be controlled with our proposed model.

Author supplied keywords

Cite

CITATION STYLE

APA

Ho, T. V., & Akagi, M. (2021). Cross-Lingual Voice Conversion with Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network. IEEE Access, 9, 47503–47515. https://doi.org/10.1109/ACCESS.2021.3063519

Cross-Lingual Voice Conversion with Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network

Abstract

Author supplied keywords

Cite

Register to see more suggestions