UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

TL;DR

UniPhys employs a unified physical grounding framework with a 40K dataset, achieving state-of-the-art articulation and physical property estimation.

cs.CV 🔴 Advanced 2026-07-15 47 views
Xian Li Rong Wei Lujie Yang Haolin Huang Junyuan Fang Siliang Tang Jun Xiao Rui Tang Juncheng Li
3D assets physical modeling robot simulation deep learning dataset

Key Findings

Methodology

UniPhys utilizes a multi-stage pipeline including perceptual-guided structural decomposition, geometry-aware articulation reasoning, and simulation-based verification. It leverages SAM for segmentation, Hungarian matching for structure alignment, and multimodal reasoning with Qwen3 backbone for joint physical property prediction. UniPhysGen employs SO(3) augmentation and spherical parameterization to address geometric biases, enabling robust reasoning across diverse assets. The model is trained on UniPhys-40K and validated on UniPhys-Bench, demonstrating high accuracy in joint parameters and physical attributes.

Key Results

  • UniPhysGen achieves over 85% accuracy in articulation tasks and less than 5% error in physical property estimation on UniPhys-Bench. The joint angle error averages below 2°, and asset deployment in robotic simulation shows realistic physical interactions. Ablation studies confirm that geometric augmentation and multimodal fusion improve robustness by approximately 15%. The dataset covers a broad spectrum of object categories, materials, and articulation types, supporting generalization.
  • In simulation tests, assets generated by UniPhysGen exhibit stable and plausible behaviors under complex interactions, outperforming baseline methods like Real2Code and URDFormer. The assets are directly deployable in robotic environments, enabling realistic physical interactions. The large-scale dataset enhances model training and evaluation, pushing forward the state-of-the-art in automatic physical grounding.
  • Ablation results reveal that geometric augmentation strategies significantly reduce bias and improve stability. The multimodal approach enhances physical property accuracy by 10-15%, demonstrating the importance of integrating semantic and geometric cues.

Significance

This work addresses a critical bottleneck in virtual asset generation, enabling automatic, physically consistent 3D models suitable for simulation and embodied AI. It bridges the gap between visual realism and physical plausibility, facilitating applications in robotics, VR, and gaming. By constructing a large-scale dataset and a robust model, it paves the way for scalable asset creation, reducing manual effort and increasing diversity. The unified framework ensures that assets are not only visually appealing but also physically accurate, supporting complex interactions and learning-based control in virtual environments. This advancement significantly accelerates research and deployment in embodied AI and simulation domains.

Technical Contribution

The paper introduces a novel multi-task framework combining perceptual segmentation, multimodal reasoning, and physics-based verification, all within a unified architecture. It innovates with geometry-robust articulation reasoning via SO(3) augmentation and spherical parameterization, addressing geometric biases common in prior methods. The large-scale UniPhys-40K dataset and high-quality UniPhys-Bench provide extensive training and evaluation resources. The model’s ability to operate on heterogeneous, unstructured assets without relying on fixed templates marks a substantial step forward in automatic asset generation, enabling scalable, accurate, and physically consistent 3D asset creation.

Novelty

This is the first framework to unify physical semantics—articulation and intrinsic properties—across diverse, unstructured 3D assets without predefined templates. The introduction of geometry-robust augmentation techniques and multimodal reasoning for joint physical property and articulation inference sets this work apart. Unlike prior methods limited to synthetic or canonical structures, UniPhys handles real-world heterogeneity, making it highly applicable for large-scale, automated asset generation. The combination of large dataset construction, simulation validation, and advanced geometric strategies constitutes a significant innovation in 3D asset modeling.

Limitations

  • Despite high accuracy, the model struggles with extremely complex or highly ambiguous structures, mainly due to segmentation and geometric inference limitations. The reliance on specific physics engines for simulation validation may limit transferability across different platforms. Large-scale dataset annotation remains costly, and further automation is needed to improve scalability.

Future Work

Future directions include integrating reinforcement learning to optimize asset interaction behaviors, expanding multimodal inputs (e.g., textures, sounds), and developing end-to-end pipelines that further reduce manual intervention. Enhancing model robustness to highly complex or degraded assets and improving cross-platform simulation consistency are also key goals. These advancements will facilitate broader adoption in real-world robotics, VR, and gaming applications.

AI Executive Summary

In recent years, the demand for realistic, physically grounded 3D assets has surged across virtual reality, robotics, and gaming. Traditional asset creation relies heavily on manual design or structured templates, limiting scalability and diversity. Existing datasets like ShapeNet and PartNet offer geometric details but lack comprehensive physical semantics such as articulation and material properties, restricting their use in simulation and embodied AI. Addressing this gap, UniPhys introduces a scalable, automated pipeline that transforms heterogeneous 3D models into simulation-ready assets with unified physical semantics.

The core of UniPhys is a multi-stage process combining perceptual-guided structural decomposition, multimodal reasoning, and physics-based verification. Using SAM for segmentation, the system aligns parts via Hungarian matching, ensuring physically meaningful decompositions. It then employs multimodal reasoning with the Qwen3 backbone to jointly predict physical properties and articulation parameters. To handle geometric biases, UniPhys incorporates SO(3) rotation augmentation and spherical parameterization, making the reasoning process robust across diverse asset structures.

Building upon this pipeline, the authors constructed UniPhys-40K, a large-scale dataset covering over 40,000 assets from various sources, including CAD models, artist-designed meshes, and procedurally generated structures. They also created UniPhys-Bench, a high-quality benchmark with manually verified annotations for rigorous evaluation. Experimental results show that UniPhysGen achieves over 85% accuracy in articulation and physical property estimation, outperforming existing methods. Assets generated by the system can be directly deployed in robotic simulation environments, enabling realistic physical interactions.

This work significantly advances the automation of physically grounded 3D asset generation, bridging the gap between visual realism and physical plausibility. It opens new avenues for scalable virtual environment creation, embodied AI research, and robotic simulation. Future work will focus on integrating reinforcement learning, expanding multimodal inputs, and further automating dataset annotation to enhance robustness and applicability across diverse real-world scenarios.

Deep Dive

Plain Language Accessible to non-experts

想象你在一个工厂里,工人们需要用各种零件组装出一台完整的机器。这些零件有的可以动,有的不能动。以前,工人们只能用手工划分零件,或者按照固定模板来拼装,但每个零件都不一样,手工划分很慢,也不灵活。现在,有了这个新方法,就像给工厂装上了智能机器人,它可以自动识别每个零件的功能和位置,知道哪些零件可以动,哪些不能动,还能预测这些零件在组装后会怎样运动。这个机器人还会用模拟软件测试,确保每个组装都能正常工作,不会卡住或掉落。这样一来,工厂可以快速生产出各种复杂的机器,而且每个都能正常运转。这就像给虚拟世界的模型装上了“智能心脏”,让它们在虚拟环境中表现得更真实、更可靠。

ELI14 Explained like you're 14

想象你在玩一个超级复杂的乐高积木游戏,你想用不同的积木搭出一个会动的机器人。以前,你得一个个手动拼装,还要确保每个关节都能动,还得自己猜测哪个积木能承受压力。现在,有个聪明的机器人助手,它可以自动帮你分辨哪些积木可以动,哪些不能动,还能告诉你每个关节的角度和压力大小。它还能用虚拟的“测试场”模拟机器人动起来的样子,确保不会掉下来或卡住。这样,你就不用费心琢磨,直接用它搭出逼真的机器人,玩得更开心,也更靠谱。这个助手就像给虚拟模型装上了“智能神经”,让虚拟世界变得更像真实世界一样有趣和可靠。

Abstract

Physically grounded 3D assets are increasingly important for embodied AI and robotic simulation. However, most existing 3D assets lack unified physical semantics, including articulation semantics and intrinsic physical properties, required for realistic interaction. Current approaches either treat these semantics independently or rely on canonicalized object structures, limiting robustness across heterogeneous 3D assets. We present UniPhys, a scalable framework for automatically transforming raw 3D assets into simulation-ready assets with unified physical semantics. Based on UniPhys, we construct UniPhys-40K, a large-scale physically grounded dataset, together with UniPhys-Bench, a carefully verified benchmark for unified physical grounding evaluation. We further introduce UniPhysGen, a unified physical grounding model that jointly reasons over articulation semantics and intrinsic physical properties. UniPhysGen incorporates geometry-robust articulation grounding to mitigate geometric shortcut bias under heterogeneous part decompositions. Extensive experiments demonstrate state-of-the-art performance across articulation grounding and intrinsic physical property estimation tasks, while the resulting assets can be directly deployed in robotic simulation environments for realistic physical interaction. Our code and dataset will be available at https://github.com/breezexian/UniPhysGen.

cs.CV