CMA-OT

Hierarchical Expert Supervision for Dance-to-Music Generation

Abstract

Dance-to-Music (D2M) generation aims to synthesize music that aligns rhythmically and stylistically with dance videos. Existing methods typically rely on sparse dance cues (e.g., rhythm and style) and supervise only the final audio output, which leads to a semantic gap and limits the learning of expressive musical representations. To address these challenges, we propose CMA-OT (Curriculum-guided Multi-scale representation learning with scale-aware Optimal Transport), a framework that leverages a pre-trained music expert to provide hierarchical supervision across multiple semantic levels. CMA-OT includes: Curriculum-guided multi-scale learning, which progressively transfers musical knowledge from coarse to fine levels for effective representation guidance. Scale-aware Fused Gromov-Wasserstein Optimal Transport (FGW-OT), which models soft correspondences between hierarchical expert representations and latent generator features. Extensive experiments on benchmark datasets show that CMA-OT achieves state-of-the-art performance in rhythmic synchronization, perceptual quality, and overall music generation.

Method Pipeline

Method Pipeline

Generation Results

Generation Comparison with SOTA Dance-to-Music Methods.

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)

D2M_GAN

CDCD

LORIS

Textual_Inv

MotionComposer

CMA-OT (Ours)