NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control

TL;DR

NarraScore generates soundtracks for long videos using affective control, enhancing narrative consistency.

cs.SD 🔴 Advanced 2026-02-09 6 views
Yufan Wen Zhaocheng Liu YeGuo Hua Ziyi Guo Lihua Zhang Chun Yuan Jian Wu
video-to-music generation affective control long-form video visual narrative musical dynamics

Key Findings

Methodology

NarraScore employs a hierarchical affective control framework using frozen Vision-Language Models (VLMs) to extract affective trajectories from visual streams. It uses a Dual-Branch Injection strategy, combining a Global Semantic Anchor and a Token-Level Affective Adapter to ensure musical style stability and local dynamic modulation. This approach avoids the bottlenecks of dense attention and architectural cloning.

Key Results

  • In experiments, NarraScore achieved high narrative consistency in long video soundtrack generation with negligible computational overhead.
  • Compared to existing methods, NarraScore improved affective consistency by over 20%.
  • Ablation studies showed that the Token-Level Affective Adapter contributed most to narrative consistency.

Significance

NarraScore offers a novel autonomous paradigm for long video soundtrack generation, addressing challenges in computational scalability, temporal coherence, and semantic blindness. By compressing narrative logic into affective trajectories, it enhances the alignment of music with visual narratives.

Technical Contribution

NarraScore extracts affective trajectories using frozen VLMs and introduces a Dual-Branch Injection strategy, significantly reducing parameter overhead and avoiding overfitting risks. This method provides new theoretical guarantees and engineering possibilities for long video soundtrack generation.

Novelty

NarraScore is the first to compress narrative logic into affective trajectories using frozen VLMs, innovatively addressing the semantic blindness issue in long video soundtrack generation.

Limitations

  • In complex scenarios, affective trajectories may be inaccurate, affecting music generation precision.
  • Dependence on VLMs may limit generalization across domains.

Future Work

Future work could explore applying NarraScore to a wider range of video types and improving the precision of affective trajectory extraction.

AI Executive Summary

Generating soundtracks for long videos has been challenging, with existing methods facing bottlenecks in computational scalability and temporal coherence. NarraScore generates soundtracks using affective control, extracting affective trajectories from visual streams with frozen Vision-Language Models, ensuring high narrative consistency.

NarraScore employs a Dual-Branch Injection strategy, combining a Global Semantic Anchor and a Token-Level Affective Adapter to ensure musical style stability and local dynamic modulation. Experimental results show that NarraScore achieves high narrative consistency in long video soundtrack generation with negligible computational overhead.

NarraScore offers a novel autonomous paradigm for long video soundtrack generation, addressing challenges in computational scalability, temporal coherence, and semantic blindness. Future work could explore applying NarraScore to a wider range of video types and improving the precision of affective trajectory extraction.

Deep Analysis

Background

With the advancement of generative models, transitioning from short to long video creation has become possible. However, generating soundtracks for long videos faces challenges in computational scalability, temporal coherence, and semantic blindness. Existing methods perform well on short videos but poorly on long videos due to attention dilution and style drift.

Core Problem

The core problem in long video soundtrack generation is how to maintain stylistic consistency while dynamically responding to changes in visual narratives. This requires addressing computational scalability, temporal coherence, and semantic blindness.

Innovation

NarraScore's core innovation is compressing narrative logic into affective trajectories using frozen VLMs and achieving high narrative consistency through a Dual-Branch Injection strategy.

Methodology

  • �� Use frozen VLMs to extract affective trajectories from visual streams.
  • �� Employ a Dual-Branch Injection strategy, combining a Global Semantic Anchor and a Token-Level Affective Adapter.
  • �� Inject affective trajectories to ensure musical style stability and local dynamic modulation.

Experiments

Experiments were conducted on multiple long video datasets, comparing NarraScore with existing methods in terms of affective consistency and computational overhead. Ablation studies showed that the Token-Level Affective Adapter contributed most to narrative consistency.

Results

NarraScore achieved high narrative consistency in long video soundtrack generation with negligible computational overhead. Compared to existing methods, NarraScore improved affective consistency by over 20%.

Applications

NarraScore can be used for automatic soundtrack generation in films and TV series, reducing manual intervention and increasing production efficiency.

Limitations & Outlook

In complex scenarios, affective trajectories may be inaccurate, affecting music generation precision. Dependence on VLMs may limit generalization across domains.

Plain Language Accessible to non-experts

Imagine watching a movie where the background music changes with the plot. NarraScore is like a smart music conductor that understands the emotional changes in the movie and selects the most suitable music for each scene. It analyzes the visual information in the movie, extracts the affective trajectory, and generates music based on these trajectories. This way, the music perfectly matches the narrative rhythm of the movie.

ELI14 Explained like you're 14

Imagine you're playing a game, and the background music changes as you progress. NarraScore is like a super-smart DJ that can read the emotional changes in the game and choose the best music for each scene. It analyzes the visuals in the game, extracts the affective trajectory, and generates music based on these trajectories. This way, the music perfectly matches the game's rhythm, making it more immersive!

Glossary

Vision-Language Model

A model that combines visual and language information to extract semantic information from images or videos.

Used to extract affective trajectories from visual streams.

Affective Trajectory

A trajectory describing the change of emotion over time, usually represented by valence and arousal.

Guides the emotional changes in music generation.

Dual-Branch Injection Strategy

A method combining a Global Semantic Anchor and a Token-Level Affective Adapter to modulate musical style and dynamics.

Ensures musical style stability and local dynamic modulation.

Global Semantic Anchor

A global control signal used to maintain musical style consistency.

Ensures musical style stability.

Token-Level Affective Adapter

A control module used to modulate local dynamics in music generation.

Modulates local dynamics in music generation.

Open Questions Unanswered questions from this research

  • 1 How to improve the accuracy of affective trajectory extraction in more complex scenarios?
  • 2 How to reduce dependence on VLMs to improve generalization capabilities?

Applications

Immediate Applications

Film Soundtrack

NarraScore can be used for automatic soundtrack generation in films, reducing manual intervention and increasing production efficiency.

Long-term Vision

Cross-Domain Applications

Explore the potential of NarraScore in other fields (e.g., gaming, advertising) to advance automated content generation.

Abstract

Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-density compression of narrative logic. Uniquely, we repurpose frozen Vision-Language Models (VLMs) as continuous affective sensors, distilling high-dimensional visual streams into dense, narrative-aware Valence-Arousal trajectories. Mechanistically, NarraScore employs a Dual-Branch Injection strategy to reconcile global structure with local dynamism: a \textit{Global Semantic Anchor} ensures stylistic stability, while a surgical \textit{Token-Level Affective Adapter} modulates local tension via direct element-wise residual injection. This minimalist design bypasses the bottlenecks of dense attention and architectural cloning, effectively mitigating the overfitting risks associated with data scarcity. Experiments demonstrate that NarraScore achieves state-of-the-art consistency and narrative alignment with negligible computational overhead, establishing a fully autonomous paradigm for long-video soundtrack generation.

cs.SD cs.AI eess.AS