On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
Proposes MM-Det leveraging LMM-generated multi-modal forgery representations, achieving 92.0% AUC on DVF for diffusion video detection.
Key Findings
Methodology
The proposed MM-Det integrates large-scale multi-modal models (LMMs) to generate multi-modal forgery representations (MMFR) by combining visual encoders like CLIP with large language models such as Llama 2. An In-and-Across Frame Attention (IAFA) mechanism captures spatial artifacts and temporal inconsistencies. Video reconstruction via VQ-VAE amplifies diffusion traces, aiding forgery detection. The framework employs end-to-end training on a newly constructed large-scale diffusion video dataset (DVF), with a dynamic fusion strategy to adaptively combine multi-modal and spatio-temporal features, enhancing generalization and robustness.
Key Results
- On the DVF dataset, MM-Det achieved an average AUC of 92.0%, outperforming the previous best (HiFi-Net at 84.3%). It maintained high accuracy across various diffusion models, with detection rates exceeding 70% in most scenarios, reaching 95% in some cases. Ablation studies confirmed that MMFR contributed a 7% performance boost, and IAFA added 5%. The model demonstrated stable performance across different resolutions and video lengths, confirming its strong generalization.
- Compared to baseline methods like F3Net and HiFi-Net, MM-Det showed significant improvements in detecting diverse forgery types, especially in complex scenes with subtle artifacts. The model's sensitivity to spatial and temporal cues enabled it to identify fake videos with minimal artifacts, even in challenging conditions.
- The results validate that multi-modal representations and attention mechanisms substantially enhance forgery detection accuracy, especially for unseen or sophisticated forgeries.
Significance
This work advances video forensic techniques by leveraging the perceptual and reasoning capabilities of LMMs, addressing limitations of prior methods that relied on single-modal features. It provides a robust, generalizable framework capable of detecting a wide range of diffusion-generated videos, thus strengthening defenses against malicious deepfake and synthetic media. The construction of the DVF dataset also supplies a valuable benchmark for future research, fostering progress in multimedia security and authenticity verification.
Technical Contribution
The study introduces a novel multi-modal forgery representation (MMFR) derived from LMMs, combined with a new In-and-Across Frame Attention (IAFA) mechanism for capturing spatial and temporal forgery cues. It innovatively employs VQ-VAE for efficient video reconstruction to amplify diffusion traces. The end-to-end training pipeline, supported by a large-scale diffusion video dataset (DVF), sets a new standard for generalizable forgery detection. These contributions significantly extend the capabilities of existing detection frameworks, enabling better handling of diverse and unseen forgery scenarios.
Novelty
This research is the first to incorporate large-scale multi-modal models into video forgery detection, creating a multi-modal feature space that captures complex semantic, spatial, and temporal cues. The IAFA mechanism uniquely balances local patch-level and global frame-level information, enhancing detection sensitivity. Unlike prior methods limited to static images or single features, this approach effectively addresses the dynamic and semantic diversity of diffusion-generated videos, marking a significant innovation.
Limitations
- Despite high accuracy, the model may still struggle with videos containing extremely subtle or heavily masked forgery traces, especially in low-quality or highly compressed videos. The reliance on large-scale annotated datasets could limit adaptability to unseen forgery types not represented during training. Computational complexity remains high, posing challenges for real-time deployment in resource-constrained environments.
Future Work
Future directions include exploring self-supervised pretraining of multi-modal models to reduce dependence on labeled data, integrating additional modalities like audio or text for multi-source forgery detection, and optimizing model architectures for real-time inference. Further research will also focus on improving robustness against adversarial attacks and extending detection capabilities to emerging generative paradigms.
AI Executive Summary
Recent advances in diffusion models have revolutionized video synthesis, enabling the creation of highly realistic fake videos with diverse semantics. However, this progress poses significant challenges for content authenticity verification, as traditional detection methods mainly target facial manipulations and lack robustness against complex, semantic-rich forgery content. Existing approaches often rely on single-modal features, such as frequency artifacts or facial inconsistencies, which are insufficient for the broad spectrum of diffusion-generated videos.
In response, this study introduces MM-Det, a novel detection framework that leverages the perceptual and reasoning strengths of large multi-modal models (LMMs). By generating a multi-modal forgery representation (MMFR) that fuses visual and textual cues, MM-Det captures subtle semantic anomalies and forgery traces. The core innovation is the In-and-Across Frame Attention (IAFA) mechanism, which aggregates local patch-level features and global frame-level cues, effectively detecting spatial artifacts and temporal inconsistencies. To amplify diffusion artifacts, the framework employs VQ-VAE-based video reconstruction, computing residual differences that highlight forgery traces.
Experimental results on the newly constructed Diffusion Video Forensics (DVF) dataset demonstrate that MM-Det achieves an average AUC of 92.0%, surpassing prior state-of-the-art methods by a significant margin. The model maintains robust performance across various diffusion models, resolutions, and video lengths, confirming its strong generalization. These findings suggest that multi-modal, attention-based detection strategies are highly effective for complex, semantic-rich forgery scenarios.
This work significantly advances the field of video forensics, providing a scalable, generalizable tool for combating increasingly sophisticated synthetic media. The construction of the DVF dataset also offers a valuable benchmark for future research. Moving forward, integrating additional modalities, optimizing computational efficiency, and exploring self-supervised learning are promising directions to further enhance detection capabilities and real-world deployment.
Deep Dive
Plain Language Accessible to non-experts
想象你在一家工厂里,工厂每天都在生产各种商品。有些商品是真品,有些是伪造的。工厂的检测员需要用各种工具和线索来判断商品的真伪。传统的方法就像只看商品的外包装或标签,容易被伪造品蒙混过去。现在,工厂引入了一台超级智能的检测机器,它不仅能看外包装,还能听到生产线的声音、分析材料,甚至用特殊的扫描仪检测商品内部的微小差异。这台机器就像本文中的多模态检测系统,它结合了多种信息源,能更准确地识别伪造品。它用的技术就像是让机器拥有“眼睛”、“耳朵”和“脑袋”,一起合作,找到那些微妙的伪造痕迹。这样一来,即使有人用最先进的技术伪造视频,也难以骗过它的“火眼金睛”。这项技术就像给电脑装上了“超级眼睛”和“超级大脑”,让它成为打击虚假内容的强大武器。
ELI14 Explained like you're 14
想象你在学校里,有一个超级厉害的老师,他不仅能看出谁在作弊,还能听出谁在说谎。这位老师用的不是普通的眼睛或耳朵,而是装了各种高科技设备,比如可以扫描作业的仪器、能听出微弱声音的麦克风,还有能分析作业内容的电脑。这些设备让老师变得无比聪明,能发现平时难以察觉的作弊行为。类似的,科学家们也在研究一种超级智能的“老师”,它可以通过分析视频中的各种细节,判断视频是否是伪造的。它不仅看视频的画面,还能分析声音、动作、甚至视频中的微小差异。这样一来,即使有人用最先进的技术伪造视频,也难以骗过它的“火眼金睛”。这项技术就像给电脑装上了“超级眼睛”和“超级大脑”,让它成为打击虚假内容的强大武器。
Abstract
Large numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identify diffusion-generated content with a diverse range of semantics. To advance the field of video forensics, we propose an innovative algorithm named Multi-Modal Detection(MM-Det) for detecting diffusion-generated videos. MM-Det utilizes the profound perceptual and comprehensive abilities of Large Multi-modal Models (LMMs) by generating a Multi-Modal Forgery Representation (MMFR) from LMM's multi-modal space, enhancing its ability to detect unseen forgery content. Besides, MM-Det leverages an In-and-Across Frame Attention (IAFA) mechanism for feature augmentation in the spatio-temporal domain. A dynamic fusion strategy helps refine forgery representations for the fusion. Moreover, we construct a comprehensive diffusion video dataset, called Diffusion Video Forensics (DVF), across a wide range of forgery videos. MM-Det achieves state-of-the-art performance in DVF, demonstrating the effectiveness of our algorithm. Both source code and DVF are available at https://github.com/SparkleXFantasy/MM-Det.