Song2Face: Synthesizing Singing Facial Animation from Audio

12Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We present Song2Face, a deep neural network capable of producing singing facial animation from an input of singing voice and singer label. The network architecture is built upon our insight that, although facial expression when singing varies between different individuals, singing voices store valuable information such as pitch, breathe, and vibrato that expressions may be attributed to. Therefore, our network consists of an encoder that extracts relevant vocal features from audio, and a regression network conditioned on a singer label that predicts control parameters for facial animation. In contrast to prior audio-driven speech animation methods which initially map audio to text-level features, we show that vocal features can be directly learned from singing voice without any explicit constraints. Our network is capable of producing movements for all parts of the face and also rotational movement of the head itself. Furthermore, stylistic differences in expression between different singers are captured via the singer label, and thus the resulting animations singing style can be manipulated at test time.

Cite

CITATION STYLE

APA

Iwase, S., Kato, T., Yamaguchi, S., Yukitaka, T., & Morishima, S. (2020). Song2Face: Synthesizing Singing Facial Animation from Audio. In SIGGRAPH Asia 2020 Technical Communications, SA 2020. Association for Computing Machinery, Inc. https://doi.org/10.1145/3410700.3425435

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free