Inference-time Physics Alignment of Video Generative Models with Latent World Models
Improved video generation physics plausibility using WMReward and VJEPA-2, achieving 62.64% in ICCV 2025 challenge.
Key Findings
Methodology
This study introduces WMReward, a method that leverages the physics prior of the latent world model VJEPA-2 as a reward signal during inference to enhance the physics plausibility of video generation. The method optimizes generation performance by searching and guiding multiple candidate denoising trajectories.
Key Results
- In the ICCV 2025 Perception Test PhysicsIQ Challenge, the WMReward method achieved a final score of 62.64%, surpassing the previous state-of-the-art by 7.42%.
- Significant improvements in physics plausibility were observed across image-conditioned, multiframe-conditioned, and text-conditioned generation settings.
- Human preference studies validated the method's effectiveness, showing an 11.4% improvement in physics plausibility.
Significance
This research demonstrates the feasibility of using latent world models to improve the physics plausibility of video generation, addressing the shortcomings of existing models in this area. The advancement is significant for academia and industry, particularly in applications like robotics and autonomous driving.
Technical Contribution
The technical contribution lies in using latent world models for inference-time alignment in video generation, providing a novel reward model to guide the generation process. This method shows superior physics understanding compared to existing vision-language model-based approaches.
Novelty
This study is the first to use the physics prior of latent world models for inference-time alignment in video generation, introducing the WMReward reward model, which offers more effective improvements in physics plausibility compared to existing methods.
Limitations
- The method relies on the accuracy of the latent world model; if the model's understanding of physical phenomena is inaccurate, it may affect results.
- Searching and guiding multiple candidate trajectories can lead to high computational costs, especially with limited resources.
Future Work
Future research could explore more efficient search strategies to reduce computational costs and apply this method to a broader range of generation tasks, such as 3D scene generation.
AI Executive Summary
Current video generation models have made significant progress in generating visual content, yet they still fall short in terms of physics plausibility. Many models lack an understanding of physical laws during pre-training, resulting in videos that do not adhere to physical common sense. To address this issue, researchers have introduced a new method called WMReward, which leverages the physics prior of the latent world model VJEPA-2 as a reward signal during inference to enhance the physics plausibility of video generation.
The WMReward method optimizes generation performance by searching and guiding multiple candidate denoising trajectories. In the ICCV 2025 Perception Test PhysicsIQ Challenge, this method achieved a final score of 62.64%, surpassing the previous state-of-the-art by 7.42%. Experimental results show significant improvements in physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings.
The significance of this method lies in demonstrating the feasibility of using latent world models to improve the physics plausibility of video generation, addressing the shortcomings of existing models in this area. This advancement is significant for academia and industry, particularly in applications like robotics and autonomous driving. Future research could explore more efficient search strategies to reduce computational costs and apply this method to a broader range of generation tasks, such as 3D scene generation.
Deep Analysis
Background
Video generation technology has seen significant advancements, particularly in generating visual content. However, these models still fall short in terms of physics plausibility, often generating videos that do not adhere to basic physical laws. This issue not only affects user experience but also limits the application of these models in fields like robotics and autonomous driving. Existing research has primarily focused on improving physics plausibility through pre-training enhancements, but inference-time alignment remains underexplored.
Core Problem
Current video generation models face challenges in generating physically plausible videos. Despite significant improvements in visual quality, the generated videos often violate basic physical laws. The core issue lies in the models' lack of deep understanding of physical phenomena during inference, leading to results that do not align with real-world physical common sense.
Innovation
The core innovation of this study is the introduction of WMReward, a method that leverages the physics prior of the latent world model VJEPA-2 as a reward signal during inference to enhance the physics plausibility of video generation. This method shows superior physics understanding compared to existing vision-language model-based approaches.
Methodology
- �� Utilize VJEPA-2 model's physics prior as a reward signal
- �� Search and guide multiple candidate denoising trajectories during inference
- �� Optimize generation performance using the WMReward model
- �� Validate the method's effectiveness in the ICCV 2025 challenge
Experiments
Experiments were conducted across multiple generation settings, including image-conditioned, multiframe-conditioned, and text-conditioned setups. Models used include MAGI-1, Sora2, and vLDM. Evaluation metrics included physics plausibility scores and human preference studies. Results showed significant improvements in physics plausibility across all settings.
Results
Experimental results showed that the WMReward method achieved a score of 62.64% in the ICCV 2025 challenge, surpassing the previous state-of-the-art by 7.42%. Additionally, human preference studies showed an 11.4% improvement in physics plausibility. These results demonstrate the significant advantage of the WMReward method in enhancing video generation physics plausibility.
Applications
This method can be directly applied to video generation tasks requiring high physics plausibility, such as robotic navigation and autonomous driving. By improving the physics plausibility of generated videos, these applications can achieve more reliable environment modeling and decision support.
Limitations & Outlook
Despite the WMReward method's outstanding performance in enhancing physics plausibility, it incurs high computational costs, especially in resource-constrained scenarios. Additionally, the method relies on the accuracy of the latent world model; if the model's understanding of physical phenomena is inaccurate, it may affect results.
Plain Language Accessible to non-experts
Imagine you're in a kitchen cooking. Existing video generation models are like a chef who only focuses on the appearance of ingredients; the food might look good but taste off. The WMReward method is like an experienced chef who considers not just the appearance but also the taste and nutrition of ingredients. By using the physics prior of latent world models, WMReward acts like a guidebook for the chef, ensuring each dish not only looks good but also meets health standards.
ELI14 Explained like you're 14
Imagine you're playing a virtual reality game where objects look real, but when you try to interact with them, they don't follow physical laws, like a ball not rolling or water not flowing. The WMReward method is like a new game patch that fixes these unrealistic physics. By using a model called VJEPA-2, this patch ensures that every object in the game moves according to real-world physics, making the game experience more realistic and fun!
Glossary
Latent World Model
A predictive model that encodes high-dimensional observations into compact latent representations and learns the transition function in this latent space to forecast future states.
Used in the paper to provide physics priors to enhance video generation physics plausibility.
VJEPA-2
A latent world model trained through self-supervised learning, achieving state-of-the-art performance on physics understanding benchmarks.
Used as the reward model in the WMReward method.
WMReward
A reward model that leverages the physics prior of latent world models to enhance the physics plausibility of video generation.
Used during inference to guide video generation models.
Denoising Trajectories
The process of gradually generating target data through denoising steps during generation.
Used in the WMReward method for candidate trajectory search.
ICCV 2025 Perception Test PhysicsIQ Challenge
A challenge evaluating the physics plausibility of video generation models.
The WMReward method achieved excellent results in this challenge.
Open Questions Unanswered questions from this research
- 1 How to improve the efficiency of the WMReward method without increasing computational costs?
- 2 What is the applicability of latent world models in different generation tasks?
- 3 How to further enhance the semantic consistency of the WMReward method?
Applications
Immediate Applications
Robotic Navigation
By improving the physics plausibility of video generation, robots can more accurately perceive and understand their environment, enhancing navigation and decision-making capabilities.
Long-term Vision
Autonomous Driving
Applying this method in autonomous driving can improve the vehicle's understanding and prediction of the environment, enhancing safety and reliability.
Abstract
State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems from suboptimal inference strategies. We therefore introduce WMReward and treat improving physics plausibility of video generation as an inference-time alignment problem. In particular, we leverage the strong physics prior of a latent world model (here, VJEPA-2) as a reward to search and steer multiple candidate denoising trajectories, enabling scaling test-time compute for better generation performance. Empirically, our approach substantially improves physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings, with validation from human preference study. Notably, in the ICCV 2025 Perception Test PhysicsIQ Challenge, we achieve a final score of 62.64%, winning first place and outperforming the previous state of the art by 7.42%. Our work demonstrates the viability of using latent world models to improve physics plausibility of video generation, beyond this specific instantiation or parameterization.