PWM-ArtGen: Part World Model for Articulated Object Generation
PWM-ArtGen uses diffusion-based joint modeling of visual dynamics and kinematic parameters, outperforming baselines with strong zero-shot generalization.
Key Findings
Methodology
PWM-ArtGen employs a diffusion-based joint model combining action diffusion and image diffusion, learning the joint distribution of visual dynamics and kinematic parameters. It uses a noise prediction network with independent diffusion timesteps for unannotated data, enhanced by a visual dynamics regularizer (VDR). The dataset comprises 19.7k high-fidelity part-level image pairs without annotations, enabling large-scale training. The model excels in static and dynamic states, demonstrating superior accuracy and zero-shot generalization on multiple datasets.
Key Results
- On ACD and PartNet-Mobility, PWM-ArtGen surpasses SOTA with over 20% improvement in dgIoU, significantly reducing geometric errors, achieving 100% articulation graph accuracy, and demonstrating robust zero-shot performance in complex scenes.
- The model effectively leverages unlabeled data via joint diffusion, reducing annotation costs and enhancing adaptability.
- VDR improves dynamic behavior fidelity, validated by qualitative and quantitative metrics, confirming the importance of visual motion cues in kinematic inference.
Significance
This work addresses fundamental challenges in single-image articulated object generation by integrating visual dynamics and kinematic reasoning within a scalable diffusion framework. It overcomes data scarcity and improves generalization, opening new avenues for robotics, virtual reality, and digital twin applications. The approach reduces reliance on costly annotations and enhances the realism of generated articulated models, marking a significant advance in 3D generative modeling.
Technical Contribution
The paper introduces a novel diffusion transformer architecture that jointly models visual and kinematic information, with independent diffusion steps enabling effective learning from unlabeled data. The integration of VDR ensures dynamic consistency, and the large-scale dataset supports robust training. The framework bridges static image understanding with dynamic 3D modeling, offering a new paradigm for scalable articulated object synthesis.
Novelty
This is the first work to leverage diffusion models for joint learning of visual dynamics and kinematic parameters from a single image, utilizing independent diffusion timesteps for unlabeled data. The incorporation of a visual dynamics regularizer and large-scale high-fidelity dataset further distinguishes this approach, enabling high-quality, zero-shot generalization to complex articulated objects, a significant leap over prior static or multi-view methods.
Limitations
- The model struggles with extreme occlusions and highly complex articulation patterns due to limited training data diversity. Computational cost remains high because of diffusion sampling steps, limiting real-time applications.
- High-quality data collection and augmentation are costly, and current datasets do not fully cover all possible articulation types or textures. Future work should focus on reducing inference latency and expanding dataset diversity.
- While generalization is strong, fine-grained control over specific joint behaviors remains challenging, requiring further refinement for industrial deployment.
Future Work
Future research will explore multi-modal data fusion, such as integrating depth or temporal information, to improve dynamic inference. Efforts will also target model acceleration, enabling real-time applications. Expanding datasets with more diverse and complex articulated objects will further enhance robustness and applicability across industries.
AI Executive Summary
PWM-ArtGen represents a significant breakthrough in single-image articulated object generation, leveraging a diffusion-based joint model that captures both visual dynamics and kinematic parameters. Traditional methods often rely on multi-view data or expensive annotations, limiting scalability and real-world applicability. In contrast, PWM-ArtGen employs a unified framework that combines action diffusion and image diffusion, enabling the model to learn the joint distribution of motion and appearance from large-scale unlabeled data.
The core innovation lies in the use of independent diffusion timesteps, which facilitate effective co-training on unlabeled datasets. This approach, complemented by a visual dynamics regularizer (VDR), ensures that generated dynamic behaviors are physically plausible and visually consistent. The authors curated a dataset of 19.7k high-fidelity part-level image pairs, supporting large-scale training and bridging the synthetic-real gap.
Experimental results demonstrate that PWM-ArtGen outperforms existing methods on multiple benchmarks, achieving over 20% improvement in key geometric metrics and perfect articulation graph accuracy in zero-shot scenarios. The model's ability to generate realistic, articulated 3D objects from a single image opens new avenues for robotics, virtual reality, and digital twin applications, where rapid, accurate, and scalable modeling is critical.
Looking ahead, the authors plan to incorporate multi-modal data, optimize inference speed, and expand datasets to cover more complex articulations. This work sets a new standard for scalable, high-fidelity articulated object generation, promising broad impact across both academia and industry.
Deep Analysis
Background
The evolution of 3D object reconstruction has transitioned from multi-view optimization to single-image generation, driven by deep learning advances. Early methods like PartNet-Mobility relied heavily on annotated datasets, which are costly and limited in diversity. Recent approaches, such as implicit surface representations and large-scale foundation models, have improved geometric fidelity but struggle with dynamic behaviors. The challenge remains to generate articulated objects that are both geometrically accurate and dynamically plausible, especially from a single image. Existing datasets like ACD and synthetic datasets like PartNet-Mobility provide valuable benchmarks but lack the scale and realism needed for robust generalization. Consequently, there is a pressing need for models that can leverage unannotated data effectively, capturing complex visual dynamics and kinematic structures in diverse real-world scenarios.
Core Problem
The core challenge in single-image articulated object generation is accurately inferring kinematic parameters and dynamic behaviors without explicit supervision. Static images lack explicit motion cues, making it difficult to determine joint types, axes, and ranges. Existing methods either depend on multi-view data or manual annotations, which are expensive and not scalable. The fundamental bottleneck is the scarcity of large-scale, high-fidelity, unannotated datasets that can support learning complex visual dynamics. This limits the ability of models to generalize to real-world objects with diverse structures and textures. Overcoming these issues requires innovative approaches that can learn from abundant unlabeled visual data while ensuring physically plausible and geometrically consistent outputs.
Innovation
This work introduces several key innovations: 1) a diffusion-based joint model that learns the distribution of visual dynamics and kinematic parameters simultaneously, 2) independent diffusion timesteps enabling effective co-training on unlabeled data, 3) a visual dynamics regularizer (VDR) that enforces dynamic consistency, and 4) a large-scale high-fidelity dataset of 19.7k part-level image pairs. Unlike prior static or multi-view methods, PWM-ArtGen captures the interplay between motion and appearance, allowing for accurate, controllable generation of articulated objects from a single image. The integration of these components creates a scalable framework that bridges the gap between synthetic and real data, facilitating zero-shot generalization and dynamic realism.
Methodology
- �� Input: single image and target part structure. • Part mask generator: uses SAM and GPT-4o to extract part masks and structural graphs. • Joint diffusion model: employs a diffusion transformer to predict kinematic parameters and visual dynamics, conditioned on static image features and part references. • Unsupervised training: utilizes independent diffusion steps for unlabeled data, with a coupled noise prediction network. • Visual regularization: applies VDR to align generated dynamics with real motion cues, enhancing physical plausibility. • Data: constructs 19.7k high-fidelity, unannotated part-level image pairs via rendering and editing, supporting large-scale training.
Experiments
The model is trained on PartNet-Mobility and evaluated on ACD and unseen PartNet-Mobility objects. Metrics include dgIoU, dcDist, dCD, and articulation accuracy. Baselines include URDFormer, NAP, and SINGAPO. Ablation studies assess the impact of VDR and unlabeled data. Hyperparameters include 200k training steps, 12-layer diffusion transformer, and AdamW optimizer. Evaluation covers static and dynamic states, with zero-shot tests on out-of-distribution objects. Results show significant improvements over baselines, especially in complex scenarios, validating the effectiveness of the joint diffusion approach.
Results
PWM-ArtGen achieves over 20% higher dgIoU and lower geometric errors compared to SOTA on both datasets. Zero-shot generalization maintains high accuracy, with 100% articulation graph correctness. Visualizations confirm realistic motion and geometry, even in occluded or textured objects. Ablation results demonstrate that removing VDR or unlabeled training reduces performance, highlighting their importance. The model also produces more coherent and physically plausible articulated behaviors, validated through quantitative metrics and qualitative assessments.
Applications
This approach can be directly applied to robotic manipulation, enabling robots to understand and generate complex articulated objects from minimal input. It also benefits virtual reality content creation, allowing rapid generation of realistic animated characters or objects. In industrial automation, it facilitates automatic modeling of machinery and components, reducing manual effort. Long-term, integrating multi-modal data like depth and temporal sequences could further enhance dynamic accuracy, supporting autonomous systems and digital twins in real-world environments.
Limitations & Outlook
Despite strong performance, the model faces challenges with highly occluded or extremely complex articulation patterns due to limited data diversity. Computational costs remain high because of the diffusion sampling process, restricting real-time deployment. Data collection and augmentation are expensive, and current datasets do not cover all possible object types or textures. Future work should focus on reducing inference latency, improving robustness in challenging scenarios, and expanding datasets to include more diverse and intricate articulations.
Plain Language Accessible to non-experts
想象你在玩一个拼装玩具,比如乐高积木。每次你只拿到一张图片,里面显示了一个复杂的机械或动物模型,但没有说明怎么拼装。PWM-ArtGen就像是一个聪明的拼装专家,它可以根据那张图片,猜出每个积木块应该放在哪里,还能知道每个关节会怎么动。它通过学习很多不同的拼装方式,知道哪些积木可以转动、伸缩,然后自己拼出一个逼真的三维模型。就像是给机器人装上了“眼睛”和“脑袋”,让它能理解和创造复杂的机械结构,甚至在没有人告诉它具体细节的情况下,也能做出合理的动作和姿势。
ELI14 Explained like you're 14
想象你在玩一个超级复杂的拼图游戏,只不过这个拼图是三维的,而且每个拼图块还能动!你只有一张图片,显示了拼图的最终样子,但没有告诉你每块应该怎么拼。PWM-ArtGen就像是一个聪明的拼图大师,它可以根据那张图片,猜出每个拼图块的形状和它们之间的关系,还能知道每个关节会怎么动。它用一种特别的方法,学习了很多拼图的秘密,比如哪些块可以转动,哪些可以伸缩。这样,即使没有说明书,它也能拼出一个逼真的3D模型,而且还能让模型动起来,就像是魔法一样!这让机器人变得更聪明,能自己理解复杂的机械结构和动作。
Abstract
The key challenge in articulated 3D object generation from a single image is accurately predicting the underlying kinematic structure. Existing methods either infer kinematic parameters directly from a static image that lacks dynamic part-level kinematic relationships, or estimate parameters from visual dynamics generated from a single image, which is prone to accumulated errors of two steps. Moreover, the limited scale and diversity of existing annotated datasets further hinder generalization to complex, real-world objects. To overcome these limitations, we propose to learn the joint distribution of visual dynamics and kinematic parameters. Recognizing that articulated objects can be formulated as dynamic systems, we propose a unified Part World Model called PWM-ArtGen. To leverage unannotated data, this model couples action diffusion and image diffusion with independent diffusion timesteps, which enables visual branch co-training. We further curate a photorealistic dataset of 19.7k part-level image pairs without kinematic annotations, to support co-training. Experiments demonstrate that PWM-ArtGen substantially outperforms existing baselines in the resting state and exhibits strong zero-shot generalization to out-of-distribution objects.