DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
DreamID-Omni framework achieves controllable human-centric audio-video generation, surpassing existing commercial models.
Key Findings
Methodology
The study introduces the DreamID-Omni framework, which uses a Symmetric Conditional Diffusion Transformer to integrate heterogeneous conditioning signals. It employs a Dual-Level Disentanglement strategy and a Multi-Task Progressive Training scheme to address identity-timbre binding failures and speaker confusion in multi-person scenarios.
Key Results
- In the R2AV task, DreamID-Omni achieved an ID-Sim. score of 0.674/0.603, surpassing the existing Wan2.6 model.
- In the RV2AV task, it scored 0.584 in AES and 14.832 in ViCLIP, outperforming VACE and HunyuanCustom.
- In the RA2V task, it achieved a Sync-C score of 6.325, showing better audio-visual synchronization.
Significance
This research provides a unified framework for audio-video generation, addressing the challenge of integrating multiple tasks. It holds significant academic and commercial value, especially in scenarios requiring precise control over character identities and voice timbres.
Technical Contribution
Technical contributions include the introduction of a Symmetric Conditional Diffusion Transformer and Dual-Level Disentanglement strategy, solving identity confusion in multi-person generation and enhancing model generalization through multi-task progressive training.
Novelty
This is the first framework to unify R2AV, RV2AV, and RA2V tasks, overcoming the limitations of isolated task processing and significantly enhancing generation flexibility and consistency.
Limitations
- In complex backgrounds with multiple characters, identity and timbre disentanglement remains challenging.
- The model requires substantial computational resources during training, which may limit practical applications.
Future Work
Future work could explore more efficient training methods to reduce computational resource requirements and further improve disentanglement in multi-character scenarios.
AI Executive Summary
Recent advancements in audio-video generation have been significant, yet existing methods typically treat human-centric audio-video generation, video editing, and audio-driven video animation as isolated tasks. The DreamID-Omni framework successfully integrates these tasks using a Symmetric Conditional Diffusion Transformer and Dual-Level Disentanglement strategy, addressing identity-timbre binding failures and speaker confusion in multi-person scenarios.
Experimental results demonstrate that DreamID-Omni achieves leading performance in video, audio, and audio-visual consistency, even surpassing some commercial models. Its multi-task progressive training scheme effectively prevents overfitting and enhances coordination across different tasks.
Nonetheless, DreamID-Omni faces challenges in disentangling identity and timbre in complex scenarios. Future research could explore more efficient training methods to reduce computational resource demands and further enhance performance in multi-character scenarios.
Deep Analysis
Background
Audio-video generation technology has made significant strides in recent years, particularly driven by foundational models. However, existing methods typically treat human-centric audio-video generation, video editing, and audio-driven video animation as isolated tasks, lacking a unified framework to integrate these tasks.
Core Problem
Existing methods struggle to achieve precise control over multiple character identities and voice timbres within a single framework, particularly in multi-person scenarios where identity-timbre binding failures and speaker confusion are common.
Innovation
The core innovation of DreamID-Omni lies in introducing a Symmetric Conditional Diffusion Transformer and Dual-Level Disentanglement strategy, achieving unified integration of multiple tasks and significantly enhancing generation flexibility and consistency.
Methodology
- �� Symmetric Conditional Diffusion Transformer: Integrates heterogeneous conditioning signals
- �� Dual-Level Disentanglement Strategy: Signal-level Syn-RoPE and semantic-level Structured Captions
- �� Multi-Task Progressive Training: Gradual training from weakly-constrained to strongly-constrained tasks
Experiments
Experiments used the IDBench-Omni benchmark, including 100 identity-timbre-caption triplets, 50 masked videos with target identity and timbre, and 50 driving audios with reference identities.
Results
In the R2AV task, DreamID-Omni achieved an ID-Sim. score of 0.674/0.603, surpassing the existing Wan2.6 model. In the RV2AV task, it scored 0.584 in AES and 14.832 in ViCLIP, outperforming VACE and HunyuanCustom.
Applications
The framework can be used in film production, virtual reality, and game development scenarios requiring precise control over character identities and voice timbres.
Limitations & Outlook
Despite significant progress, DreamID-Omni faces challenges in disentangling identity and timbre in complex backgrounds with multiple characters, and the model requires substantial computational resources during training.
Plain Language Accessible to non-experts
Imagine a concert where each band member has their own instrument and style. DreamID-Omni is like a super conductor, able to coordinate multiple band members simultaneously, ensuring they play a harmonious and beautiful symphony on the same stage. It ensures each member's unique style is not confused while adjusting their performance as needed.
ELI14 Explained like you're 14
Imagine you're playing a super cool game where you can create your own characters, not just choosing their appearance but also their voice! DreamID-Omni is like a magic tool in the game, allowing you to easily switch characters' looks and voices, letting them interact seamlessly in the game world. Isn't that awesome?
Glossary
Symmetric Conditional Diffusion Transformer
A transformer that integrates heterogeneous conditioning signals for unified task generation.
Used to integrate reference images, voice timbres, source videos, and driving audio.
Dual-Level Disentanglement
A method to solve identity confusion in multi-person scenarios through signal and semantic level disentanglement.
Ensures accurate identity and timbre binding in multi-person generation.
Synchronized RoPE
A mechanism to bind reference identities with corresponding voice timbres at the signal level.
Addresses identity-timbre binding issues in multi-person scenarios.
Structured Captions
Establishes explicit attribute-subject mappings through anchor tokens and fine-grained descriptions.
Resolves semantic-level confusion in multi-person scenarios.
Multi-Task Progressive Training
A training strategy that gradually introduces tasks with varying constraints to enhance model generalization.
Coordinates learning objectives across different tasks.
Open Questions Unanswered questions from this research
- 1 How to improve identity and timbre disentanglement in complex backgrounds?
- 2 How to reduce computational resource demands to enhance practical feasibility?
Applications
Immediate Applications
Film Production
In film production, precise control over character identities and voice timbres is needed, and DreamID-Omni can help achieve more natural character performances.
Long-term Vision
Virtual Reality
Achieving more realistic character interactions in virtual reality, DreamID-Omni can significantly enhance user experience.
Abstract
Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and audio-driven video animation (RA2V) as isolated objectives. Furthermore, achieving precise, disentangled control over multiple character identities and voice timbres within a single framework remains an open challenge. In this paper, we propose DreamID-Omni, a unified framework for controllable human-centric audio-video generation. Specifically, we design a Symmetric Conditional Diffusion Transformer that integrates heterogeneous conditioning signals via a symmetric conditional injection scheme. To resolve the pervasive identity-timbre binding failures and speaker confusion in multi-person scenarios, we introduce a Dual-Level Disentanglement strategy: Synchronized RoPE at the signal level to ensure rigid attention-space binding, and Structured Captions at the semantic level to establish explicit attribute-subject mappings. Furthermore, we devise a Multi-Task Progressive Training scheme that leverages weakly-constrained generative priors to regularize strongly-constrained tasks, preventing overfitting and harmonizing disparate objectives. Extensive experiments demonstrate that DreamID-Omni achieves comprehensive state-of-the-art performance across video, audio, and audio-visual consistency, even outperforming leading proprietary commercial models. We will release our code to bridge the gap between academic research and commercial-grade applications.