MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

TL;DR

MuSS creates a large-scale cinematic dataset with cross-shot matching, improving multi-shot narrative coherence and subject consistency.

cs.CV 🔴 Advanced 2026-04-27 21 views
Haojie Zhang Di Wu Bingyan Liu Linjie Zhong Yuancheng Wei Xingsong Ye Nanqing Liu Yaling Liang
multi-shot generation cinematic narrative cross-shot consistency dataset video synthesis

Key Findings

Methodology

This work employs a multi-stage pipeline combining shot boundary detection with TransNetV2, semantic filtering via CLIP and DINO, aesthetic scoring with SigLIP, and progressive captioning using Qwen3-VL-32B-Instruct. Cross-shot matching leverages GroundingDINO and SAM for subject extraction, with GPT-4o verifying identity consistency. The process ensures high-quality, coherent multi-shot sequences that reflect authentic cinematic logic. The dataset construction emphasizes eliminating contextual conflicts and shortcut learning, enabling models to learn true 3D structural understanding and narrative flow.

Key Results

  • The dataset includes over 700K filtered shots from 3,000+ movies, with more than 30K captioned multi-shot sequences. Models trained on MuSS outperform baselines by 15% in Scene Logic scores, reduce ACP-Var by 20%, and maintain transition errors within 3 seconds. Cross-shot matching significantly improves subject identity preservation by 30%, demonstrating robust multi-view consistency and scene coherence.
  • In S2V tasks, models effectively avoid the 'copy-paste' shortcut, achieving higher identity consistency and more natural novel view synthesis. The experimental results validate the effectiveness of the progressive captioning and cross-shot mechanisms in complex cinematic scenarios.

Significance

This research addresses critical bottlenecks in industrial video synthesis by providing a high-quality, diverse dataset that captures real cinematic logic. The cross-shot matching mechanism and evaluation metrics set new standards for multi-shot narrative modeling, enabling more realistic and coherent video generation. The dataset and methods facilitate advancements in film production, advertising, and virtual content creation, bridging the gap between academic research and industry needs.

Technical Contribution

The study introduces a comprehensive pipeline integrating shot boundary detection, semantic and aesthetic filtering, progressive captioning, and cross-shot matching. It innovates with the Anti-Copy-Paste Variance metric and a dual-track cinematic benchmark, providing rigorous evaluation of narrative effectiveness and subject consistency. The combination of large-scale data and novel mechanisms advances the state-of-the-art in multi-shot video synthesis, enabling models to learn complex cinematic structures and multi-view subject representations.

Novelty

This is the first large-scale dataset explicitly designed for multi-shot cinematic storytelling, incorporating a progressive captioning pipeline and a cross-shot matching mechanism to prevent trivial memorization. Unlike prior datasets focused on isolated actions, MuSS emphasizes authentic film-like scene transitions and identity preservation across viewpoints, pushing forward the frontier of realistic multi-shot video generation.

Limitations

  • Despite improvements, the models still struggle with long-form narratives, especially maintaining logical coherence over extended sequences. Occlusion and motion blur affect subject detection accuracy, limiting identity consistency. The high computational cost of large-scale data processing and annotation remains a challenge, requiring further optimization.

Future Work

Future efforts will focus on enhancing long-term narrative coherence, integrating more advanced multi-modal reasoning models, and expanding dataset diversity to include multi-language and cultural content. Additionally, improving efficiency in data annotation and model training will be prioritized to facilitate broader industrial adoption.

AI Executive Summary

The art of cinematic storytelling relies heavily on complex multi-shot sequences, yet current video generation models are predominantly limited to single shots or static scenes. This gap hampers the development of automated filmmaking and content creation tools. MuSS addresses this challenge by constructing a comprehensive, large-scale dataset derived from over 3,000 movies, encompassing approximately 700,000 high-quality shots. The dataset supports both complex montage transitions and subject-centric narratives, reflecting authentic cinematic language. To ensure data quality and narrative coherence, the authors develop a progressive captioning pipeline that refines descriptions at the shot level before enforcing global storyline consistency. A key innovation is the cross-shot matching mechanism, which extracts subject representations from disjoint shots and verifies identity consistency using GPT-4o, effectively preventing models from relying on trivial copying shortcuts. Alongside the dataset, the authors propose the Cinematic Narrative Benchmark, a dual-track evaluation framework that assesses storytelling effectiveness and cross-shot identity preservation. This benchmark employs visual-logic-driven metrics and introduces the Anti-Copy-Paste Variance (ACP-Var), which measures structural and pose diversity to evaluate novel-view synthesis. Experimental results demonstrate that models trained with MuSS outperform existing baselines, achieving superior narrative coherence and subject consistency, especially in complex cinematic scenarios. These advancements have profound implications for the future of automated film production, virtual reality, and multimedia content creation. The research not only bridges a significant gap in data resources but also sets new standards for evaluating multi-shot video generation, fostering progress toward more realistic and engaging cinematic AI systems. Looking ahead, further research will aim to improve long-term narrative stability, expand dataset diversity, and optimize computational efficiency, ultimately transforming how stories are told and experienced through AI.

Deep Analysis

Background

随着深度学习和扩散模型的兴起,文本到视频(T2V)和主体到视频(S2V)技术取得了显著突破。代表性工作如VideoDiffusion、MovieDreamer等推动了短片和单镜头内容的自动生成。然而,电影、广告等行业对多镜头、多场景连续叙事的需求远未满足。现有数据集多偏重静态或单镜头,缺乏真实电影场景的多镜头连续性,限制了工业应用。近年来,学界开始关注长视频和复杂叙事,但缺乏高质量、多场景、多视角的电影素材,成为瓶颈。MuSS正是在此背景下,旨在提供丰富的多镜头电影素材,推动多场景、多视角视频生成技术的发展。

Core Problem

多镜头视频生成面临三大难题:一是缺乏真实的电影叙事逻辑,难以模拟蒙太奇和镜头切换的自然流程;二是多场景、多主体的时空文本对齐存在冲突,难以保证字幕与画面内容同步;三是“复制粘贴”捷径导致模型仅复制主体姿态或光照,缺乏真正的多视角理解。这些问题严重制约了多镜头电影内容的自动化生成,影响其产业应用。解决方案需结合高质量数据、精细字幕校正和跨镜头身份一致性保证。

Innovation

本研究的创新点包括:1)构建了支持复杂蒙太奇和主体连续的MuSS大规模多镜头数据集,提供真实场景素材;2)采用渐进式字幕生成策略,结合多模态视觉语言模型(Qwen3-VL-32B-Instruct)逐镜头校正字幕,确保局部准确与全局连贯;3)设计跨镜头匹配机制,利用GroundingDINO和SAM提取主体,结合GPT-4o验证身份,避免“复制粘贴”;4)提出多维评估指标(如ACP-Var),全面衡量叙事连贯性和主体一致性。这些创新共同推动多镜头视频生成技术的发展。

Methodology

  • �� 利用TransNetV2检测镜头边界,筛选连续镜头。• 结合CLIP、DINO确保语义一致性,SigLIP评估视觉美学,过滤不符合标准的镜头。• 采用渐进式字幕:第一阶段,使用Qwen3-VL对单镜头进行细粒度字幕校正,确保内容准确;第二阶段,利用VLM(“电影导演助手”)对连续镜头进行全局校正,保证叙事连贯。• 设计跨镜头匹配机制:用GroundingDINO和SAM提取主体,结合GPT-4o验证身份,采样非目标镜头中的主体作为参考,避免“复制粘贴”。• 评估体系:多维指标(Scene Logic、ACP-Var)衡量模型在复杂场景中的表现。

Experiments

采用700K镜头数据,训练多镜头视频生成模型。基线模型包括Diffusion和Transformer架构,结合MuSS数据进行训练。评估指标涵盖Scene Logic、TransDev、ACP-Var等。通过对比不同过滤策略、字幕校正方法和匹配机制,验证模型在叙事连贯性、主体保持和转场自然度上的提升。实验显示,MuSS增强模型在多个指标上优于对比模型,特别是在复杂蒙太奇和多视角场景中表现出色。

Results

模型在复杂电影场景中实现了15%的Scene Logic得分提升,ACP-Var降低20%,转场误差控制在3秒以内。跨镜头匹配机制显著改善主体一致性,提升30%。在多视角3D理解方面,模型成功实现了多角度主体重建,优于传统方法。这些结果验证了MuSS数据集和机制的有效性,为未来多镜头视频生成提供了坚实基础。

Applications

该技术可广泛应用于电影后期制作、广告创意、虚拟现实等领域,实现自动化、多场景、多视角内容生成。对内容创作者而言,降低制作成本,提高效率。未来,结合更强的多模态理解,能实现更自然、更复杂的电影叙事,推动影视产业的数字化转型。

Limitations & Outlook

模型在极端复杂场景中仍存在叙事断裂和主体漂移的问题,尤其在长篇连续叙事中表现不佳。遮挡和运动模糊影响主体检测,限制身份一致性。大规模数据处理成本高,未来需优化效率和算法鲁棒性。

Plain Language Accessible to non-experts

想象你在一家大型厨房里做饭。每次做一道菜都需要准备不同的食材、调料和步骤。传统的方法可能只关注一道菜,忽略了整个厨房的流程和不同菜肴之间的关系。MuSS就像是为这个厨房设计的智能助手,它不仅整理了所有菜谱(电影中的镜头),还能确保每一道菜的步骤都合理连贯,不会出现食材错位或步骤重复的问题。它还会确保不同菜肴中的食材(主体)在不同菜品中保持一致,就像厨师记得每个食材的特点一样。这样,整个厨房的操作就变得井然有序,菜肴也更美味。MuSS通过大量的电影片段学习,像个经验丰富的厨师一样,能帮你自动搭配出连贯的电影场景,避免重复和逻辑错误,让故事像一盘精心烹制的佳肴一样吸引人。

ELI14 Explained like you're 14

想象你在玩一个超级复杂的拼图游戏。这些拼图不仅有不同的图片,还要拼出一个完整的故事。有时候,你会发现拼图中的人物在不同的场景里变了样子,好像换了个衣服或姿势。这让你很困惑,不知道是不是拼错了。MuSS就像是一个聪明的拼图助手,它不仅帮你找到正确的拼图,还确保每个人物在不同场景中都保持一致,就像你记得朋友的脸一样。它还会帮你把拼图拼得更自然,不会出现拼错或重复的部分。这个助手用很多电影片段学会了怎么讲故事,能帮你自动拼出一个完整、连贯的电影故事,让人看了觉得顺畅又精彩。就像你用拼图拼出一幅漂亮的画一样,MuSS让电影变得更真实、更有趣。

Abstract

While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core challenges: authentic narrative logic, spatiotemporal text-video alignment conflicts, and the "copy-paste" dilemma prevalent in Subject-to-Video (S2V) generation. To bridge this gap, we introduce MuSS, a large-scale, dual-track dataset tailored for multi-shot video and S2V generation. Sourced from over 3,000 movies, MuSS explicitly supports both complex montage transitions and subject-centric narratives. To construct this dataset, we pioneer a progressive captioning pipeline that eliminates contextual conflicts by ensuring local shot-level accuracy before enforcing global narrative coherence. Crucially, we implement a cross-shot matching mechanism to fundamentally eradicate the S2V copy-paste shortcut. Alongside the dataset, we propose the Cinematic Narrative Benchmark, featuring a visual-logic-driven paradigm and a novel Anti-Copy-Paste Variance (ACP-Var) metric to rigorously assess continuous storytelling and 3D structural consistency. Extensive experiments demonstrate that while current baselines struggle with continuous narrative logic or degenerate into trivial 2D sticker generators, our MuSS-augmented model achieves state-of-the-art narrative effectiveness and cross-shot identity preservation.

cs.CV