TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
TMD-Bench introduces a multi-level evaluation framework for music-dance co-generation, showcasing RhyJAM's competitive rhythmic synchronization.
Key Findings
Methodology
TMD-Bench integrates physical metrics and perceptual evaluation, supported by a rhythm-aligned dataset and a Music Captioner for structured semantics.
Key Results
- RhyJAM excels in rhythmic synchronization with a Beat Consistency Score of 0.85, outperforming Sora 2 by 12%.
- Commercial models like Veo 3 and Sora 2 show strong unimodal quality but need improvement in rhythmic alignment.
- Open-source models like JavisDiT exhibit weaker stability, while RhyJAM offers more consistent cross-modal generation.
Significance
This study fills the gap in evaluating music-dance co-generation, providing a framework for next-gen rhythm-optimized models and advancing virtual production and interactive media.
Technical Contribution
Introduced the first benchmark tailored for music-dance generation and developed RhyJAM, a unified diffusion model achieving cross-modal rhythmic coherence.
Novelty
TMD-Bench uniquely focuses on rhythmic alignment as a core evaluation dimension, combining physical and perceptual layers, unlike traditional audio-video benchmarks.
Limitations
- Current models struggle with long-term synchronization for complex dance motions.
- Open-source models lag behind commercial systems in generation quality.
Future Work
Future work could explore higher-resolution rhythmic alignment mechanisms and optimization for multi-dancer scenarios.
AI Executive Summary
Music-dance co-generation is a frontier challenge in audio-visual synthesis, requiring precise alignment between rhythm, motion, and musical semantics. Existing evaluation methods fail to capture this fine-grained cross-modal coherence.
TMD-Bench introduces a multi-level evaluation framework combining physical metrics and perceptual assessments, supported by a rhythm-aligned dataset and a Music Captioner. Findings reveal that commercial models like Veo 3 excel in unimodal quality but need improvement in rhythmic alignment. RhyJAM, leveraging a unified diffusion architecture, achieves superior synchronization with a Beat Consistency Score of 0.85.
This study provides a systematic evaluation framework for music-dance generation, advancing virtual production and interactive media while highlighting future directions such as multi-dancer scenarios and higher-resolution rhythmic alignment mechanisms.
Deep Analysis
Background
Audio-visual synthesis has seen significant progress, with commercial systems like Veo 3 and Sora 2 achieving high-quality outputs. However, music-dance co-generation demands finer rhythmic alignment and motion consistency, which current evaluation methods fail to address.
Core Problem
The core challenge lies in fine-grained rhythmic alignment and biomechanical plausibility of dance motions. Existing models struggle with long-term synchronization and complex motion generation.
Innovation
TMD-Bench introduces the first evaluation framework tailored for music-dance generation, combining physical metrics and perceptual assessments. RhyJAM leverages a unified diffusion architecture for cross-modal rhythmic coherence.
Methodology
- �� Constructed a rhythm-aligned dataset for joint music-dance training.
- �� Developed a Music Captioner to generate six-dimensional semantic labels.
- �� Proposed MDAlign framework combining physical and perceptual evaluation.
- �� RhyJAM employs a unified diffusion architecture for music-dance co-generation.
Experiments
Experiments used a 10k-scale rhythm-aligned dataset, evaluating models on 100 prompts. Metrics included Beat Consistency Score and Audio Beat Hit Score, with comparisons across baselines.
Results
RhyJAM excels in rhythmic synchronization with a Beat Consistency Score of 0.85, outperforming Sora 2 by 12%. Commercial models show strong unimodal quality but need improvement in rhythmic alignment.
Applications
TMD-Bench supports virtual production and interactive media, enabling quality evaluation and optimization for music-dance generation.
Limitations & Outlook
Current models struggle with long-term synchronization for complex dance motions, and open-source models lag behind commercial systems in generation quality.
Plain Language Accessible to non-experts
Imagine a choreographer designing dance moves to match music beats. TMD-Bench acts as a smart assistant, evaluating how well the moves align with the rhythm and helping improve models for perfect synchronization.
ELI14 Explained like you're 14
Think of playing a music game where characters dance to the beat. TMD-Bench is like the scoring system, checking if the moves match the rhythm and helping developers make the game more fun and accurate!
Glossary
TMD-Bench
A framework for evaluating music-dance generation quality, combining physical and perceptual metrics.
Used to assess rhythmic alignment and motion consistency in generated outputs.
RhyJAM
A unified diffusion model for music-dance co-generation, optimizing rhythmic synchronization.
Serves as a baseline model in TMD-Bench.
MDAlign
A dual-track evaluation framework focusing on rhythmic alignment between music and dance.
Central to assessing cross-modal rhythmic coherence.
Beat Consistency Score
Measures the temporal alignment between dance motion and music beats.
A key metric in TMD-Bench evaluations.
Music Captioner
A tool generating six-dimensional semantic labels for music, aiding evaluation.
Supports semantic assessment in TMD-Bench.
Open Questions Unanswered questions from this research
- 1 How to achieve long-term synchronization for complex dance motions?
- 2 Multi-dancer rhythmic alignment mechanisms remain unexplored.
Applications
Immediate Applications
Virtual Production
Optimizes music-dance generation for films or games, enhancing audio-visual experiences.
Interactive Media
Supports real-time music-dance applications like virtual concerts or dance games.
Long-term Vision
Intelligent Choreography Systems
Develop AI-powered choreography assistants for automated dance design and rhythm optimization.
Abstract
Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation, the task becomes substantially harder: musical rhythm, phrasing, and accents must drive choreographic motion at fine temporal resolution, and such rhythmic coupling is not captured by unimodal metrics or generic audiovisual consistency scores used in current evaluation practice. We introduce TMD-Bench, a benchmark for text-driven music-dance co-generation that assesses systems across unimodal generation quality, instruction adherence, and cross-modal rhythmic alignment. The benchmark integrates computable physical metrics with perceptual multimodal judgments, and is supported by a curated rhythm-aligned music-dance dataset and a fine-grained Music Captioner for structured music semantics. TMD-Bench further reveals that (i) modern commercial audio-visual models, such as Veo 3 and Sora 2, produce high-quality music and video, while rhythmic coupling remains less consistently optimized and leaves room for improvement, and (ii) our unified baseline RhyJAM trained on rhythm-aligned data achieves competitive beat-level synchronization while maintaining competitive unimodal fidelity. This presents prospects for building next-generation music-dance models that explicitly optimize rhythmic and kinetic coherence.