Taming Diffusion Models for Music-Driven Conducting Motion Generation

N/ACitations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

Generating the motion of orchestral conductors from a given piece of symphony music is a challenging task since it requires a model to learn semantic music features and capture the underlying distribution of real conducting motion. Prior works have applied Generative Adversarial Networks (GAN) to this task, but the promising diffusion model, which recently showed its advantages in terms of both training stability and output quality, has not been exploited in this context. This paper presents Diffusion-Conductor, a novel DDIMbased approach for music-driven conducting motion generation, which integrates the diffusion model to a two-stage learning framework. We further propose a random masking strategy to improve the feature robustness, and use a pair of geometric loss functions to impose additional regularizations and increase motion diversity. We also design several novel metrics, including Fréchet Gesture Distance (FGD) and Beat Consistency Score (BC) for a more comprehensive evaluation of the generated motion. Experimental results demonstrate the advantages of our model. The code is released at https://github.com/viiika/Diffusion-Conductor.

Cite

CITATION STYLE

APA

Zhao, Z., Bai, J., Chen, D., Wang, D., & Pan, Y. (2023). Taming Diffusion Models for Music-Driven Conducting Motion Generation. In Proceedings of the Inaugural 2023 Summer Symposium Series 2023 (pp. 40–44). AAAI Press. https://doi.org/10.1609/aaaiss.v1i1.27474

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free