Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Code-as-World represents physical worlds as executable code, excelling on QuantiPhy.
Key Findings
Methodology
The study introduces Code-as-World, representing physical worlds as executable code. The methodology includes expressing physical composition, dynamic evolution, and visual appearance. An agentic discovery loop inspired by abductive reasoning iteratively optimizes executable world hypotheses.
Key Results
- Code-as-World-VL outperforms Gemini-3.1-Flash on QuantiPhy, showing significant performance improvement.
- In physical reasoning tasks, Code-as-World-VL-27B surpasses the 9B variant and all baselines.
- Verified executable worlds provide scalable physical supervision, significantly enhancing quantitative physical reasoning.
Significance
This research provides a scalable foundation for physical intelligence, addressing the lack of explicit physical mechanism representation in existing vision-language models. Executable world representations enable better understanding and reasoning of physical phenomena.
Technical Contribution
Technical contributions include a novel method for physical world representation, combining code executability with explicit physical mechanisms, offering new engineering possibilities and theoretical guarantees.
Novelty
Code-as-World is the first to represent physical worlds as executable code, offering more precise physical mechanism modeling compared to traditional visual or language representations.
Limitations
- In complex scenarios, the model may fail to capture all physical details.
- High dependency on high-quality input data.
- Further optimization of computational efficiency is needed.
Future Work
Future work includes extending to more complex physical scenarios, improving computational efficiency, and exploring broader application domains.
AI Executive Summary
Physical understanding and reasoning rely on forming compact and generalizable world representations. While modern vision-language models can recognize and explain various physical events, they often lack explicit representations of underlying mechanisms such as object states, physical parameters, and dynamics. Code-as-World represents physical worlds through executable code, providing a compact, quantitatively grounded, and controllable abstraction of the physical world. Inspired by abductive reasoning, an agentic discovery loop iteratively optimizes executable world hypotheses from multimodal observations. Experiments show Code-as-World-VL achieves state-of-the-art performance on QuantiPhy, surpassing leading proprietary models, demonstrating the potential of executable world representations as a scalable foundation for physical intelligence.
The core of Code-as-World lies in representing physical world composition, dynamic evolution, and visual appearance through code. This approach retains the mechanistic structure needed for physical reasoning while abstracting away incidental details of individual observations. Through the agentic discovery loop, the model can recover physical mechanisms from incomplete observations, supporting physical data generation, quantitative supervision, and broader downstream physical reasoning.
In applications, verified executable worlds provide scalable physical supervision for quantitative physical reasoning. Code-as-World-VL significantly improves quantitative physical reasoning capabilities, outperforming larger models like Gemini-3.1-Flash. These results indicate that executable world representations can provide scalable supervision for grounding visual models in understanding physical mechanisms. Future work will include extending to more complex physical scenarios, improving computational efficiency, and exploring broader application domains.
Deep Analysis
Background
Physical understanding is a hallmark of intelligence. Humans can reason in unfamiliar physical situations without merely memorizing individual observations. Modern vision-language models excel at describing the physical world but lack explicit representations of underlying mechanisms. Code-as-World offers a new method for representing the physical world through executable code, enabling better understanding and reasoning of physical phenomena.
Core Problem
Existing vision-language models can recognize and explain various physical events but often lack explicit representations of underlying mechanisms such as object states, physical parameters, and dynamics. This limits their application in physical reasoning.
Innovation
Code-as-World represents physical world composition, dynamic evolution, and visual appearance through executable code. This approach retains the mechanistic structure needed for physical reasoning while abstracting away incidental details of individual observations. Through the agentic discovery loop, the model can recover physical mechanisms from incomplete observations.
Methodology
- �� Physical Composition: Describes objects in the world and their physical properties.
- �� Dynamic Evolution: Describes initial object states and their temporal changes.
- �� Visual Appearance: Describes how the physical world is observed and presented.
- �� Agentic Discovery Loop: Iteratively optimizes executable world hypotheses through propose, instantiate, execute, render, and verify cycles.
Experiments
Experiments were conducted on the QuantiPhy dataset using the Code-as-World-VL model for quantitative physical reasoning. Comparisons were made with baseline models like Gemini-3.1-Flash to evaluate performance improvements in physical reasoning tasks.
Results
Code-as-World-VL achieves state-of-the-art performance on QuantiPhy, surpassing larger models like Gemini-3.1-Flash. Verified executable worlds provide scalable physical supervision for quantitative physical reasoning.
Applications
Code-as-World can be used to train vision-language models for quantitative physical reasoning, applicable in scenarios requiring understanding and reasoning of physical mechanisms, such as autonomous driving and robotics.
Limitations & Outlook
The model may fail to capture all physical details in complex scenarios and is highly dependent on high-quality input data. Future work will focus on improving computational efficiency and extending to more complex physical scenarios.
Plain Language Accessible to non-experts
Imagine a kitchen, where Code-as-World acts like an executable recipe. Each ingredient (object) has specific attributes, like weight and shape. The recipe (code) describes how ingredients interact and change. With this recipe, we can predict the final appearance and taste of a dish without trying it each time. This way, we can quickly adjust and optimize the recipe to achieve the desired outcome without wasting ingredients.
ELI14 Explained like you're 14
Imagine you're playing a game where you need to build a virtual world. Code-as-World is like a magic tool in your hands, helping you create a realistic world. You can use it to set the properties of objects, like size and weight, and then watch how they interact. This way, you can make smarter decisions in the game, like how to make a ball roll faster. Isn't that cool?
Glossary
Executable World Representation
Represents the physical world's composition, dynamic evolution, and visual appearance through code.
Used to represent the structure and mechanisms of the physical world.
Agentic Discovery Loop
A method that iteratively optimizes executable world hypotheses through propose, instantiate, execute, render, and verify cycles.
Used to iteratively optimize physical world representations.
Physical Supervision
Supervision for quantitative physical reasoning provided by verified executable worlds.
Used to train vision-language models for physical reasoning.
QuantiPhy
A dataset for evaluating quantitative physical reasoning capabilities.
Used to test the performance of the Code-as-World-VL model.
Vision-Language Model
Models capable of processing both visual and language information.
Used to recognize and explain physical events.
Open Questions Unanswered questions from this research
- 1 How to improve the model's ability to capture physical details in complex scenarios?
- 2 How to reduce dependency on high-quality input data?
- 3 How to further optimize computational efficiency?
Applications
Immediate Applications
Autonomous Driving
Enhance decision-making capabilities of autonomous systems by understanding physical mechanisms.
Robotics
Improve robotic operations in complex environments.
Long-term Vision
Smart Homes
Enhance interaction capabilities of smart home devices through physical reasoning.
Abstract
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and governing dynamics-needed for reliably reasoning how the world evolves and responds to interventions. In this work, we introduce Code-as-World, a paradigm that represents physical worlds through executable world representations. By expressing physical composition, dynamic evolution, and visual appearance as executable code, Code-as-World provides a compact, quantitatively grounded, and controllable abstraction of the physical world. To construct such representations from multimodal observations, such as natural-language descriptions or real-world videos, we develop an agentic discovery loop inspired by abductive reasoning, where an agent proposes, executes, renders, verifies, and iteratively refines executable world hypotheses. As a concrete application, we use verified executable worlds to provide scalable physical supervision for training vision-language models on quantitative physical reasoning. Experiments show that Code-as-World-VL achieves state-of-the-art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.