From Generation to Simulation: How Far Are World Models from Being True Simulators?
Using a capability-based framework, the study assesses how close generative world models are to being true simulators, highlighting gaps in physical laws and state feedback.
Key Findings
Methodology
This paper employs a capability-oriented evaluation framework based on eight core abilities of traditional simulators: asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics. It systematically reviews 200 papers from 2018 to 2026, analyzing each work’s coverage across these capabilities. The approach combines qualitative content analysis and quantitative scoring, focusing on how each technical route—latent dynamics, video generation, joint embedding—addresses these abilities. Special emphasis is placed on the presence or absence of runtime interfaces for state querying, revealing significant gaps, especially in state feedback. The methodology integrates specific algorithms (e.g., DreamerV3, Diffusion Models, V-JEPA) and experimental data to map progress and shortcomings, providing a comprehensive capability gap assessment.
Key Results
- Current generative world models excel in interaction and controllability within specific scenarios but fall short in physical law enforcement, structured state feedback, and long-horizon stability. Asset construction and physics engine capabilities are achieved in only about 19% and 17% of papers, respectively, while state feedback interfaces are rarely exposed, with only 6 papers providing runtime entity state queries.
- Latent-dynamics models like DreamerV3 demonstrate strong long-term prediction and control, yet hallucination and geometric drift remain issues. Video generation models produce high-fidelity visuals but lack physical verifiability. Joint embedding models excel in efficiency but struggle with complex physical laws, indicating a gap between visual realism and physical correctness.
- Overall, progress is notable in interaction and stability, but fundamental gaps in physical consistency, structured feedback, and long-term reproducibility hinder the models from becoming true simulators. The lack of standardized interfaces and evaluation metrics further limits practical deployment.
Significance
This capability-based assessment provides a systematic understanding of the current limitations of generative world models in replacing traditional simulators. It highlights critical gaps in physical law adherence and state feedback, guiding future research toward formal physics integration, unified interfaces, and long-term stability. The framework enables researchers and industry practitioners to quantify progress, set clear goals, and accelerate the development of trustworthy, general-purpose simulation environments. Such advancements are essential for deploying these models in robotics, autonomous driving, and virtual reality, where physical fidelity and reproducibility are paramount.
Technical Contribution
The paper introduces an external, standardized capability assessment framework based on classical simulator attributes, enabling a quantitative comparison of diverse generative models. It systematically maps 200 papers onto these capabilities, revealing specific weaknesses such as the absence of runtime state interfaces and limited physical law enforcement. The study emphasizes the importance of state feedback as a cross-route capability and proposes six future research directions, including formal physics, unified action interfaces, and long-horizon stability, providing a clear roadmap for advancing from generation to true simulation.
Novelty
This work is the first comprehensive, capability-based evaluation of generative world models against traditional simulators, moving beyond architecture-centric surveys. It uniquely quantifies the gaps in physical law adherence and state feedback, offering a concrete, measurable framework. The emphasis on external capabilities as an evaluation metric provides a novel perspective, enabling cross-route comparison and targeted improvements, especially in critical areas like physical consistency and long-term stability.
Limitations
- The assessment relies on literature content and reported interfaces, lacking direct validation of runtime capabilities and real-time performance, which may overestimate practical readiness.
- Quantitative metrics for physical correctness and long-term stability are still nascent, requiring further standardization and benchmarking.
- Most models focus on visual fidelity and efficiency, with limited exploration of complex multi-body physics and real-world dynamics, constraining their applicability as full simulators.
Future Work
Future efforts should focus on formalizing physics models, developing standardized, unified action and feedback interfaces, and establishing benchmarks for long-term stability. Integrating multi-route approaches—combining latent dynamics, video, and embeddings—may yield more robust, physically consistent simulators. Additionally, creating industry-grade evaluation tools and datasets will facilitate broader adoption and validation, accelerating the transition from generative models to reliable, general-purpose simulators.
AI Executive Summary
The rapid evolution of generative models, driven by advances in diffusion techniques, autoregressive architectures, and joint embedding methods, has sparked optimism about their potential to replace traditional simulation engines. These models can generate highly realistic images, videos, and structured states, enabling applications in gaming, robotics, and autonomous driving. However, despite impressive visual and interactive capabilities, they fall short of fulfilling the core functions of traditional simulators—namely, strict adherence to physical laws, structured state feedback, and reproducible long-term evolution.
This study systematically evaluates 200 papers from 2018 to 2026, employing a capability-based framework rooted in the eight fundamental abilities of classical simulators. The analysis reveals that while controllability, interaction, and stability have seen significant progress, asset construction and physical engine capabilities remain underdeveloped, with less than 20% coverage. Most critically, state feedback interfaces—essential for real-time control and verification—are rarely implemented, with only 6 papers providing runtime entity state queries.
The core challenge lies in hallucination and geometric drift, stemming from models learning conditional distributions rather than invariant physical laws. Latent-dynamics models like DreamerV3 demonstrate strong control and long-term prediction, but physical correctness issues persist. Video models excel visually but lack physical verifiability. The gap between visual realism and physical fidelity underscores the need for formal physics integration, unified interfaces, and long-horizon stability.
Looking ahead, the paper advocates six research directions: formalized physics, a unified action interface, first-class state feedback, long-horizon stability, downstream utility evaluation, and cross-route hybridization. These efforts aim to bridge the gap from generation to true simulation, enabling models that are not only visually impressive but also physically reliable and reproducible over extended periods. Such advancements will catalyze the deployment of trustworthy virtual environments in robotics, autonomous systems, and beyond, transforming how we simulate, plan, and learn in complex worlds.
Deep Dive
Plain Language Accessible to non-experts
想象你在一家工厂工作,工厂里有很多机器和流程,每天都在生产不同的产品。传统的模拟器就像是工厂的操作手册,告诉你每个机器怎么运转,确保每个步骤都符合规则。而现在的生成模型就像用电脑画出一幅工厂的动画,能让你看到未来的生产线会变成什么样,但它有时候会出现奇怪的地方,比如机器突然消失或者变形。这是因为它只学会了图片和视频的样子,没有真正理解工厂的物理规则。要让电脑动画变得像真实的工厂一样可靠,就需要让它学会工厂的真正运作方式,包括机器的运动、材料的流动和产品的质量控制。这就像是让动画不仅好看,还能像真实一样工作。未来的目标是让这些电脑动画变得更真实、更稳定,能像真实工厂一样长时间正常运转,这样我们就可以用它们来测试新机器、训练机器人,甚至模拟各种复杂的生产场景。
ELI14 Explained like you're 14
想象你在玩一个超级逼真的模拟游戏,但这个游戏不仅能让你看到漂亮的画面,还能像真实世界一样遵守物理规则,比如球会弹跳、车会转弯。现在的技术已经可以画出很酷的动画和场景,但它们有时候会出现奇怪的事情,比如物体突然消失或者穿墙。这是因为这些模型还没有真正理解物理的规律,只是学会了模仿图片的样子。就像你画画时知道怎么画一个球,但不知道为什么球会弹跳一样。科学家们希望让这些模型不仅能画出漂亮的场景,还能像真实世界一样遵守物理定律,能告诉我们“这个球会弹多高”、“这个车会转多快”。这样,它们就能帮助我们设计更好的机器人、自动驾驶汽车,甚至模拟未来的世界。虽然还在努力,但未来这些模型会变得越来越聪明,能帮我们解决很多现实中的难题,就像拥有一个会遵守所有规则的虚拟世界助手一样。
Abstract
With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments. Yet the remaining distance from generation to simulation lacks a systematic assessment. We present a capability-based study using an external yardstick: eight capabilities of a traditional simulator, namely asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics. We trace three main technical routes--latent dynamics, video generation, and joint-embedding prediction--and map exactly 200 representative works published from 2018 to June 2026 onto these capabilities. Our analysis shows that world models have achieved functional substitution in interaction and controllability for specific scenarios, but remain short of traditional simulators in formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution. State feedback is the most neglected cross-route shortcoming: only 6 of 163 implementation papers expose a runtime interface for querying entity states or physical parameters. We identify six research directions: formalized physics, a unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization. Project page: https://github.com/AtongWang/world-model-simulators