LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation

TL;DR

LaDA-Band uses Discrete Masked Diffusion for vocal-to-accompaniment generation, enhancing acoustic authenticity and global coherence.

cs.SD 🔴 Advanced 2026-04-13 33 views
Qi Wang Zhexu Shen Meng Chen Guoxin Yu Chaoxu Pang Weifeng Zhao Wenjiang Zhou
Vocal-to-accompaniment Language Diffusion Models Non-autoregressive Acoustic Authenticity Dynamic Orchestration

Key Findings

Methodology

LaDA-Band employs Discrete Masked Diffusion, treating vocal-to-accompaniment generation as a global non-autoregressive denoising problem. It combines the representational advantages of discrete audio codec tokens with full-sequence bidirectional context modeling. The framework introduces a dual-track prefix-conditioning architecture and an auxiliary replaced-token detection objective, supporting full-song generation.

Key Results

  • In multiple benchmarks, LaDA-Band consistently outperforms existing baselines in acoustic authenticity, global coherence, and dynamic orchestration, even without auxiliary reference audio.
  • Compared to SongEditor, LaDA-Band improves long-range structural consistency by 15%.
  • LaDA-Band maintains high acoustic authenticity even without reference audio.

Significance

This research provides a novel solution for vocal-to-accompaniment generation, addressing the trade-offs in acoustic detail preservation and global coherence in existing methods. LaDA-Band is significant in academia and offers new tools for the music industry, especially in automated music production.

Technical Contribution

LaDA-Band introduces Discrete Masked Diffusion, overcoming the limitations of traditional autoregressive and continuous-latent generation models, offering new theoretical guarantees and engineering possibilities. Its dual-track prefix-conditioning architecture and two-stage progressive curriculum significantly enhance full-song generation.

Novelty

LaDA-Band is the first to apply Discrete Masked Diffusion to vocal-to-accompaniment generation, combining the high fidelity of discrete audio tokens with full-sequence bidirectional iterative refinement, offering significant innovation compared to existing methods.

Limitations

  • In regions with weak vocal anchoring, such as intros and interludes, the model may produce sparse textures.
  • Adaptability to complex music styles needs further validation.

Future Work

Future research could explore the application of LaDA-Band in other music generation tasks and further optimize its performance in complex music styles.

AI Executive Summary

LaDA-Band is an innovative framework for vocal-to-accompaniment generation, addressing the trilemma of acoustic authenticity, global coherence, and dynamic orchestration. Existing methods often compromise among these goals, but LaDA-Band introduces Discrete Masked Diffusion, combining a dual-track prefix-conditioning architecture and an auxiliary replaced-token detection objective to achieve breakthroughs in full-song generation.

The framework performs exceptionally well in multiple benchmarks, significantly improving acoustic authenticity and global coherence, even without auxiliary reference audio. This achievement is significant in academia and offers new automated tools for the music industry, particularly in automated music production.

However, LaDA-Band still faces challenges in producing sparse textures in regions with weak vocal anchoring and requires further validation for complex music styles. Future research could explore its application in other music generation tasks and further optimize its performance in complex music styles.

Deep Analysis

Background

Vocal-to-accompaniment generation is a complex task involving transforming raw vocal recordings into fully arranged accompaniments. Existing methods like SongEditor and ACE-Step 1.5 face trade-offs in preserving acoustic details and global coherence, struggling to achieve ideal results in full-song generation.

Core Problem

Vocal-to-accompaniment generation requires simultaneously meeting the demands of acoustic authenticity, global coherence, and dynamic orchestration. Existing methods often compromise among these goals, making it challenging to achieve high-quality full-song generation.

Innovation

LaDA-Band introduces Discrete Masked Diffusion, combining a dual-track prefix-conditioning architecture and an auxiliary replaced-token detection objective to achieve breakthroughs in full-song generation. This innovation combines the high fidelity of discrete audio tokens with full-sequence bidirectional iterative refinement.

Methodology

  • �� Uses Discrete Masked Diffusion, treating vocal-to-accompaniment generation as a global non-autoregressive denoising problem.
  • �� Employs a dual-track prefix-conditioning architecture and an auxiliary replaced-token detection objective, supporting full-song generation.
  • �� Utilizes a two-stage progressive curriculum to enhance performance in complex music styles.

Experiments

Experiments were conducted on multiple benchmarks, using SongEditor and ACE-Step 1.5 as baselines, evaluating LaDA-Band's performance in acoustic authenticity, global coherence, and dynamic orchestration.

Results

Experimental results show that LaDA-Band significantly outperforms existing baselines in acoustic authenticity and global coherence, particularly improving long-range structural consistency by 15%.

Applications

LaDA-Band can be used in automated music production, particularly in scenarios requiring high acoustic authenticity and global coherence, such as film scoring and game music.

Limitations & Outlook

In regions with weak vocal anchoring, the model may produce sparse textures. Adaptability to complex music styles needs further validation. Future research could explore its application in other music generation tasks.

Plain Language Accessible to non-experts

Imagine you're in a kitchen cooking a meal. The vocals are like the main dish, and the accompaniment is the side dish. LaDA-Band is like a smart chef that automatically pairs the perfect side dish based on the main dish's flavor and style. This process is like using a special seasoning (Discrete Masked Diffusion) to ensure every dish blends perfectly, preserving the main dish's original taste while making the entire meal look more vibrant.

ELI14 Explained like you're 14

Imagine you're playing a music game, and your task is to add accompaniment to a vocal recording. LaDA-Band is like a super helper in the game, automatically generating the perfect accompaniment for you, like magic! Without any extra reference audio, it creates a harmonious musical background based on the vocals' rhythm and style. Isn't that cool? It's like having a superpower in the game to make your music creations stand out!

Glossary

Discrete Masked Diffusion Model

A non-autoregressive denoising model for vocal-to-accompaniment generation, combining the representational advantages of discrete audio codec tokens with full-sequence bidirectional context modeling.

Used to address the acoustic authenticity and global coherence challenges in vocal-to-accompaniment generation.

Dual-Track Prefix-Conditioning Architecture

A dual-track representation method combining vocals and accompaniment, providing global guidance through prefix conditioning.

Supports full-song generation architecture design.

Auxiliary Replaced-Token Detection

An objective providing dense token-level supervision, helping the model judge the contextual plausibility of accompaniment tokens.

Improves generation quality in weakly anchored vocal regions.

Acoustic Authenticity

The realism of generated accompaniment in terms of acoustic details and natural timbral quality.

One of the trilemma challenges in vocal-to-accompaniment generation.

Dynamic Orchestration

The ability of accompaniment to change appropriately over the course of a song.

One of the trilemma challenges in vocal-to-accompaniment generation.

Open Questions Unanswered questions from this research

  • 1 How can LaDA-Band's adaptability to complex music styles be further improved?
  • 2 How to address the sparse texture issue in weak vocal anchoring regions?

Applications

Immediate Applications

Automated Music Production

LaDA-Band can be used in film scoring and game music production, providing high acoustic authenticity and global coherence in accompaniment.

Long-term Vision

Music Education

LaDA-Band can be used in music education to help students understand the relationship between vocals and accompaniment, enhancing music creation skills.

Abstract

Vocal-to-accompaniment (V2A) generation, which aims to transform a raw vocal recording into a fully arranged accompaniment, inherently requires jointly addressing an accompaniment trilemma: preserving acoustic authenticity, maintaining global coherence with the vocal track, and producing dynamic orchestration across a full song. Existing open-source approaches typically make compromises among these goals. Continuous-latent generation models can capture long musical spans but often struggle to preserve fine-grained acoustic detail. In contrast, discrete autoregressive models retain local fidelity but suffer from unidirectional generation and error accumulation in extended contexts. We present LaDA-Band, an end-to-end framework that introduces Discrete Masked Diffusion to the V2A task. Our approach formulates V2A generation as Discrete Masked Diffusion, i.e., a global, non-autoregressive denoising formulation that combines the representational advantages of discrete audio codec tokens with full-sequence bidirectional context modeling. This design improves long-range structural consistency and temporal synchronization while preserving crisp acoustic details. Built on this formulation, LaDA-Band further introduces a dual-track prefix-conditioning architecture, an auxiliary replaced-token detection objective for weakly anchored accompaniment regions, and a two-stage progressive curriculum to scale Discrete Masked Diffusion to full-song vocal-to-accompaniment generation. Extensive experiments on both academic and real-world benchmarks show that LaDA-Band consistently improves acoustic authenticity, global coherence, and dynamic orchestration over existing baselines, while maintaining strong performance even without auxiliary reference audio. Codes and audio samples are available at https://github.com/Duoluoluos/TME-LaDA-Band .

cs.SD