Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

TL;DR

Wh0 framework uses generative video world models to create a 50k human-hand manipulation dataset, boosting VLA model zero-shot success to 38.9%.

cs.RO 🔴 Advanced 2026-06-21 34 views
Yangtao Chen Zixuan Chen Peiyang Wang Yong-Lu Li Jing Huo Jieqi Shi Yang Gao
generative models human-robot interaction robotic manipulation dataset zero-shot learning

Key Findings

Methodology

The Wh0 framework utilizes generative video world models to create the WM-H dataset, consisting of 50k episodes of human-object interaction videos conditioned on language, objects, and scenes. These videos are converted into robot-trainable supervision through hand motion reconstruction and visual editing, co-trained with limited real robot data to adapt pretrained VLA models to dexterous manipulation.

Key Results

  • Wh0 improves zero-shot success on 18 real-world dexterous manipulation tasks from 8.3% to 38.9%.
  • Ablation studies show scalable generation and scene/embodiment alignment are key performance drivers.
  • The WM-H dataset provides diverse human-hand interaction videos, significantly enhancing VLA model generalization.

Significance

The Wh0 framework provides scalable and controllable human-hand manipulation data sources necessary for dexterous manipulation. It addresses the trade-off between scale and scene/embodiment alignment in existing data sources, offering new data support for deploying dexterous manipulation models.

Technical Contribution

Wh0 is the first to use generative video world models to produce large-scale human-hand manipulation data. Its core contribution lies in converting generated videos into robot-trainable supervision and co-training with limited real robot data to enhance VLA model manipulation capabilities.

Novelty

Wh0 is the first framework to use generative video world models for large-scale human-hand manipulation data generation. Unlike existing methods, Wh0 emphasizes both data generation scale and scene/embodiment alignment to improve model deployment capabilities.

Limitations

  • The quality of generated videos may affect model training, especially in long-horizon videos where inconsistencies may arise.
  • The accuracy of hand reconstruction is limited by the quality of video generation, potentially introducing noise in supervision data.
  • Dependence on strong pretraining limits Wh0's applicability without pretrained models.

Future Work

Future research could explore Wh0's application in bimanual manipulation, tool use, and longer-horizon tasks. Improving video generation quality and hand reconstruction accuracy are also important directions.

AI Executive Summary

The Wh0 framework utilizes generative video world models to provide large-scale, controllable human-hand manipulation data sources, addressing the trade-off between scale and scene/embodiment alignment in existing data sources. By generating 50k episodes of human-object interaction videos conditioned on language, objects, and scenes, Wh0 converts these videos into robot-trainable supervision and co-trains with limited real robot data to adapt to dexterous manipulation.

In 18 real-world dexterous manipulation tasks, Wh0 improves zero-shot success on unseen tasks from 8.3% to 38.9%. Ablation studies show that scalable generation and scene/embodiment alignment are key drivers of performance gains. The WM-H dataset provides diverse human-hand interaction videos, significantly enhancing the generalization of VLA models.

While Wh0 makes significant advances in generating dexterous manipulation data, the quality of generated videos and the accuracy of hand reconstruction still need improvement. Additionally, Wh0's dependence on strong pretraining limits its applicability without pretrained models. Future research could explore Wh0's application in bimanual manipulation, tool use, and longer-horizon tasks.

Deep Analysis

Background

Dexterous manipulation requires generalization across objects, scenes, and tasks. Existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well-aligned with robot deployment but expensive to collect; simulation is scalable but limited by the sim-to-real gap; and real egocentric videos scale effectively but remain misaligned with robot deployment.

Core Problem

Existing data sources face a trade-off between scale and scene/embodiment alignment, limiting the generalization capabilities of dexterous manipulation models. Providing scalable and deployment-aligned data sources without relying on extensive human data collection is a significant challenge.

Innovation

The Wh0 framework uses generative video world models to create large-scale human-hand manipulation data for the first time. It converts generated videos into robot-trainable supervision through hand motion reconstruction and visual editing, co-training with limited real robot data to enhance VLA model manipulation capabilities.

Methodology

  • �� Use generative video world models to create the WM-H dataset with 50k episodes of human-object interaction videos conditioned on language, objects, and scenes.
  • �� Convert generated videos into robot-trainable supervision through hand motion reconstruction and visual editing.
  • �� Co-train with limited real robot data to adapt to dexterous manipulation.

Experiments

Experiments were conducted on a Unitree G1 humanoid robot to evaluate Wh0's performance on 18 real-world dexterous manipulation tasks. The experiments used the WM-H dataset and 400 real robot demonstration data for training, comparing different pretraining sources and adaptation data strategies.

Results

Wh0 improves zero-shot success on 18 real-world dexterous manipulation tasks from 8.3% to 38.9%. Ablation studies show scalable generation and scene/embodiment alignment are key performance drivers.

Applications

The Wh0 framework can be used for training dexterous manipulation models, particularly in scenarios requiring large-scale and deployment-aligned data sources. Its generated video data can enhance model generalization capabilities.

Limitations & Outlook

The quality of generated videos may affect model training, especially in long-horizon videos where inconsistencies may arise. The accuracy of hand reconstruction is limited by the quality of video generation, potentially introducing noise in supervision data.

Plain Language Accessible to non-experts

Imagine you're in a kitchen, needing to pick up various tools and ingredients with your hands. Wh0 acts like a virtual kitchen assistant, generating a vast amount of videos on how to manipulate these tools and ingredients with your hands. These videos not only show how to pick up and place items but also how to operate in different kitchen environments. By watching these videos, a robot learns how to work in a real kitchen, much like having a virtual internship before entering the real kitchen.

ELI14 Explained like you're 14

Imagine you're playing a super cool VR game where you can pick up all sorts of things with your hands, like cups, books, or even robot arms! Wh0 is like a super smart game assistant that generates lots of videos on how to manipulate these things with your hands. Through these videos, the robot learns how to complete these tasks in the real world, just as flexibly as you do in the game! Isn't that awesome?

Glossary

Generative Video World Model

A model used to generate video data in virtual environments, capable of producing diverse videos based on input conditions.

Wh0 uses generative video world models to create the WM-H dataset.

VLA Model

Vision-Language-Action model that integrates visual and language information for robotic control.

Wh0 enhances VLA model manipulation capabilities through generated data.

Hand Motion Reconstruction

Extracting 3D hand motion information from videos for training robotic models.

Wh0 uses hand motion reconstruction to convert videos into trainable data.

Zero-Shot Learning

The ability to learn and infer on tasks or data that have not been seen before.

Wh0 significantly improves the zero-shot learning capability of VLA models.

Ablation Study

Analyzing the impact of removing certain parts of a model on overall performance.

Ablation studies show that scalable generation and scene/embodiment alignment are key performance drivers.

Open Questions Unanswered questions from this research

  • 1 How to further improve the quality of generated videos, especially in long-horizon tasks.
  • 2 How to fully utilize Wh0-generated data without pretrained models.

Applications

Immediate Applications

Robotic Manipulation Training

Wh0-generated data can be used to train dexterous manipulation robots, enhancing their operational capabilities in real environments.

Long-term Vision

Intelligent Assistant Development

Wh0's technology can be used to develop more intelligent virtual assistants to help people complete tasks in complex environments.

Abstract

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with robot deployment but expensive to collect; simulation is scalable but limited by the sim-to-real gap; and real egocentric videos scale effectively but remain misaligned with robot deployment. We propose Wh0, a framework that uses generative video world models as scalable and controllable sources of egocentric human-hand manipulation data to unlock the manipulation capabilities of pretrained dexterous VLA models. Conditioned on language, objects, and scenes, Wh0 uses a generative world model to produce WM-H, a 50k-episode dataset of egocentric human-object interaction videos. Wh0 then converts the generated videos into robot-trainable supervision through hand motion reconstruction and visual editing. Co-trained with a limited amount of real robot data, WM-H adapts pretrained VLA models to dexterous manipulation deployment. Across 18 real-world dexterous manipulation tasks, compared with a model post-trained only on robot data, Wh0 improves zero-shot success on unseen tasks from 8.3% to 38.9%. Ablation studies further show that scalable generation and scene/embodiment alignment are key drivers of performance gains. Videos and open-source code can be found on our project website: https://chenyt31.github.io/wh0.github.io/.

cs.RO