Abstract
Engaging embodied conversational agents need to generate expressive behavior in order to be believable in socializing interactions. We present a system that can generate spontaneous speech with supporting lip movements. The neural conversational TTS voice is trained on a multi-style speech corpus that has been prosodically tagged (pitch and speaking rate) and transcribed (including tokens for breathing, fillers and laughter). We introduce a speech animation algorithm where articulatory effort can be adjusted. The facial animation is driven by time-stamped phonemes and prominence estimates from the synthesised speech waveform to modulate the lip- and jaw movements accordingly. In objective evaluations we show that the system is able to generate speech and facial animation that vary in articulation effort. In subjective evaluations we compare our conversational TTS system’s capability to deliver jokes with a commercial TTS. Both system succeeded equally good.
Author supplied keywords
Cite
CITATION STYLE
Gustafson, J., Székely, É., & Beskow, J. (2023). Generation of speech and facial animation with controllable articulatory effort for amusing conversational characters. In Proceedings of the 23rd ACM International Conference on Intelligent Virtual Agents, IVA 2023. Association for Computing Machinery, Inc. https://doi.org/10.1145/3570945.3607289
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.