I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models
I2VShield uses generative adversarial attacks to protect privacy, reducing computational costs and enhancing DiT model defenses.
Key Findings
Methodology
I2VShield is a generative adversarial attack framework specifically designed for Diffusion Transformer (DiT)-based image-to-video (I2V) models. It includes a text-adaptive perturbation generation framework and a Multimodal Attention Disruption (MAD) attack. The text-adaptive perturbation generation framework integrates adversarial learning to reduce computational overhead while maintaining visual imperceptibility. The MAD attack exploits inherent vulnerabilities of DiT models, maximizing the deviation of internal attention features.
Key Results
- On the CogVideoX-5B model, I2VShield reduced scores for Subject Consistency, Background Consistency, and Motion Smoothness by 1.5%, 1.0%, and 2.2% respectively, indicating significant disruption of spatiotemporal coherence in generated videos.
- Compared to PhotoGuard, I2VShield significantly reduced VRAM consumption and computational costs across all evaluated DiT models, demonstrating its efficiency and practical scalability.
- In the Gemini-3.1-flash-lite evaluation, I2VShield excelled in prompt consistency and motion plausibility, particularly on the OpenSora-V2-11B model.
Significance
I2VShield holds significant implications for academia and industry. It addresses the high computational resource demands of existing methods while enhancing defense effectiveness against I2V models through an innovative generative adversarial attack framework. This research provides new insights for future privacy protection technologies, especially in the context of rapid advancements in generative models, with substantial potential for safeguarding personal privacy and public safety.
Technical Contribution
I2VShield makes notable technical contributions. Unlike existing gradient-based adversarial attack methods, it achieves perturbation generation through a single forward pass using a generative adversarial network, significantly reducing computational costs. Additionally, the MAD attack enhances interference with DiT models by disrupting multimodal attention mechanisms. These innovations offer new theoretical guarantees and engineering possibilities for generative adversarial attacks.
Novelty
I2VShield is the first generative adversarial attack framework specifically tailored for DiT models. Its novelty lies in combining text-adaptive perturbation generation with multimodal attention disruption, significantly enhancing defense effectiveness against I2V models while reducing computational resource demands.
Limitations
- I2VShield requires white-box access to the target model during training, which may not be practical in some application scenarios.
- The method may still face computational challenges when handling high-resolution videos, despite significantly reducing computational costs.
Future Work
Future research could explore applying I2VShield in black-box environments and further optimizing its performance in high-resolution video generation. Additionally, investigating how this approach can be applied to other types of generative models is a promising direction.
AI Executive Summary
With the rapid development of generative artificial intelligence, image-to-video (I2V) generation models excel at producing realistic videos but also pose privacy and security issues. Existing detection methods are mostly passive and cannot defend before content generation. I2VShield provides an efficient proactive defense framework through generative adversarial attacks, specifically targeting Diffusion Transformer (DiT)-based I2V models. This method combines text-adaptive perturbation generation and multimodal attention disruption, significantly reducing computational costs while enhancing interference with generated videos. Experimental results show that I2VShield performs excellently across multiple datasets and mainstream DiT models, particularly in disrupting spatiotemporal coherence. However, the method requires white-box access to the target model during training, which may not be practical in some scenarios. Future research could explore applying I2VShield in black-box environments and further optimizing its performance in high-resolution video generation.
Deep Analysis
Background
With the rapid development of generative artificial intelligence, image-to-video (I2V) generation models excel at producing realistic videos. Early video generation methods mainly relied on Generative Adversarial Networks (GANs) and autoregressive models but faced challenges in long-term consistency and high-resolution synthesis. The emergence of diffusion models improved generation quality and training stability, while Diffusion Transformers (DiT) as scalable alternatives further enhanced physical realism and temporal consistency.
Core Problem
Despite the excellent performance of I2V models in generating realistic videos, their misuse poses privacy and security issues. Existing detection methods are mostly passive and cannot defend before content generation. Traditional gradient-based adversarial attack methods require high computational resources, making them difficult to deploy in everyday applications.
Innovation
I2VShield provides an efficient proactive defense framework through generative adversarial attacks, specifically targeting Diffusion Transformer (DiT)-based I2V models. Its innovations include combining text-adaptive perturbation generation with multimodal attention disruption, significantly reducing computational costs while enhancing interference with generated videos.
Methodology
- �� Text-Adaptive Perturbation Generation Framework: Integrates adversarial learning to reduce computational overhead while maintaining visual imperceptibility.
- �� Multimodal Attention Disruption (MAD) Attack: Exploits inherent vulnerabilities of DiT models, maximizing the deviation of internal attention features.
- �� Single Forward Pass: Eliminates iterative gradient computation during deployment, significantly reducing computational requirements.
Experiments
Experiments were conducted on the UCF101 and CelebV-Text datasets, evaluating I2VShield's defensive capability against CogVideoX-5B, OpenSora-V2-11B, and Wan2.1-14B models. Tools like VBench and Gemini-3.1-flash-lite were used for evaluation, focusing on spatiotemporal coherence and visual quality.
Results
Experimental results show that I2VShield performs excellently across multiple datasets and mainstream DiT models, particularly in disrupting spatiotemporal coherence. Compared to PhotoGuard, I2VShield significantly reduced VRAM consumption and computational costs, demonstrating its efficiency and practical scalability.
Applications
I2VShield can be used to protect personal privacy by preventing unauthorized video generation. Its low computational cost makes it suitable for resource-constrained environments such as mobile devices and edge computing.
Limitations & Outlook
Despite its excellent performance in many aspects, I2VShield requires white-box access to the target model during training, which may not be practical in some scenarios. Additionally, the method may still face computational challenges when handling high-resolution videos.
Plain Language Accessible to non-experts
Imagine you're in a kitchen cooking. You have a recipe (text prompt) and some ingredients (images). Traditional methods require you to try repeatedly to make a good dish (generate video). I2VShield is like a smart assistant that provides you with the best cooking plan (generative adversarial attack) based on the recipe and ingredients, allowing you to quickly make delicious dishes (efficient defense). This assistant not only saves you time but also ensures your dishes aren't easily copied by others (privacy protection).
ELI14 Explained like you're 14
Hey there! Imagine you're playing a game and you have a super shield (I2VShield) that protects you from bad guys. This shield is super smart and can adjust its defense strategy based on the enemy's attack, so you can easily defeat them. The coolest part is that this shield doesn't need you to spend a lot of time upgrading it; it's already strong enough to use whenever you need it. Isn't that awesome?
Glossary
Generative Adversarial Network
A deep learning model for generating data, consisting of a generator and a discriminator.
Used for generative adversarial attacks to protect privacy.
Diffusion Transformer
A generative model combining diffusion models and transformer architecture, excelling in generating high-quality videos.
The type of model specifically targeted by I2VShield.
Multimodal Attention
An attention mechanism combining multiple input modalities (e.g., image and text).
Used to disrupt spatiotemporal coherence in generated videos.
Text-Adaptive Perturbation
Perturbations generated based on text prompts to interfere with generative models.
A core component of I2VShield.
White-box Access
Complete access to a model's architecture, parameters, and gradients.
Required by I2VShield during the training phase.
Open Questions Unanswered questions from this research
- 1 How can I2VShield be applied in black-box environments? Current methods require white-box access, limiting application scenarios.
- 2 How can I2VShield's performance in high-resolution video generation be further optimized? Current methods still face computational challenges.
Applications
Immediate Applications
Privacy Protection
I2VShield can be used to protect personal images from unauthorized video generation, suitable for social media and online platforms.
Long-term Vision
Security of Generative Models
By enhancing defense capabilities against generative models, I2VShield is expected to improve the overall security of generative models in the future.
Abstract
The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has been made in detecting AI-generated videos, proactive defenses against I2V models remain underexplored. In particular, current proactive defenses against I2V models predominantly rely on gradient-based adversarial attacks, which require defenders to possess GPUs with substantial memory resources (VRAM) to generate adversarial examples. To address this issue, we propose I2VShield, a privacy protection method based on generative adversarial attacks tailored to Diffusion Transformer (DiT)-based I2V models. The proposed method primarily consists of two components: (1) a text-adaptive perturbation generation framework integrating adversarial learning to mitigate computational overhead while maintaining visual imperceptibility; and (2) an untargeted Multimodal Attention Disruption (MAD) attack that exploits the inherent vulnerabilities of DiT-based I2V models, maximizing the deviation of the internal attention features from their clean states. Extensive experiments demonstrate that our approach achieves highly competitive protection performance across various datasets and mainstream DiT-based I2V models, particularly in disrupting spatiotemporal coherence, while substantially reducing computational costs.