DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

TL;DR

DexTacWAM integrates visual and tactile data, achieving an average score of 70.6, significantly outperforming the baseline of 38.0.

cs.RO 🔴 Advanced 2026-09-22 4 views
Haoran Yuan Zekai Wang Boning Shao Haoran Lu Trevor Darrell Ismini Lourentzou Wei Zhan
Dexterous Manipulation Visuo-Tactile Learning World-Action Models Tactile Sensing Robot Learning

Key Findings

Methodology

DexTacWAM encodes each fingertip's tactile information independently, aggregates these features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile modeling. This approach was tested on six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform.

Key Results

  • DexTacWAM achieved the highest score on all six tasks, averaging 70.6 compared to the strongest baseline of 38.0.
  • Removing tactile world modeling reduced the four-task mean score from 74.7 to 26.6, highlighting the importance of modeling contact evolution.
  • Continual vision-to-touch learning was achieved with approximately 100 demonstrations per task, maintaining visual prediction quality within 0.5 dB.

Significance

DexTacWAM demonstrates how pretrained video models can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner. This approach provides new insights for contact-aware control in dexterous manipulation, addressing the limitations of traditional vision-centric models.

Technical Contribution

By introducing a finger- and pose-aware tactile compressor, DexTacWAM significantly reduces computational costs, achieving 2.26x faster training and 1.29x faster inference. The model treats tactile signals as part of the world state rather than just auxiliary inputs.

Novelty

DexTacWAM is the first to extend pretrained video priors to distributed multi-finger contact dynamics, proposing a new framework for joint visuo-tactile modeling that addresses the localization issue of tactile signals in previous studies.

Limitations

  • DexTacWAM's performance is limited by the scarcity of tactile data, as tactile data collection is costly.
  • The model's robustness in highly dynamic environments needs further validation.

Future Work

Future research could explore more efficient tactile data collection methods and validate DexTacWAM's performance in more complex manipulation tasks.

AI Executive Summary

Dexterous manipulation relies on contact dynamics, which are often only partially observable from vision. Existing World-Action Models (WAMs) are largely vision-centric and cannot directly model these contact dynamics. DexTacWAM encodes each fingertip's tactile information independently, aggregates these features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile modeling.

In six contact-rich dexterous manipulation tasks, DexTacWAM achieved the highest score on every task, averaging 70.6 compared to the strongest baseline of 38.0. Removing tactile world modeling reduced the four-task mean score from 74.7 to 26.6, highlighting the importance of modeling contact evolution. Continual vision-to-touch learning was achieved with approximately 100 demonstrations per task, maintaining visual prediction quality within 0.5 dB.

DexTacWAM demonstrates how pretrained video models can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner. This approach provides new insights for contact-aware control in dexterous manipulation, addressing the limitations of traditional vision-centric models. Future research could explore more efficient tactile data collection methods and validate DexTacWAM's performance in more complex manipulation tasks.

Deep Analysis

Background

Dexterous manipulation is crucial in robotics, especially for tasks requiring fine control. Traditional vision models have limitations in handling contact dynamics, as these dynamics are often not fully captured by vision. Recently, World-Action Models (WAMs) have shown potential in robot learning by combining video world modeling with action generation. However, these models are largely vision-centric and cannot directly model contact dynamics.

Core Problem

The core problem in dexterous manipulation is effectively modeling and utilizing contact dynamics. These dynamics are crucial for task success but are often only partially observable from vision. Existing vision models cannot directly model these dynamics, leading to poor performance in contact-rich tasks.

Innovation

DexTacWAM's core innovation is treating tactile signals as part of the world state rather than just auxiliary inputs. By introducing a finger- and pose-aware tactile compressor, DexTacWAM achieves joint visuo-tactile modeling in a data- and compute-efficient manner.

Methodology

  • �� Independently encode tactile information from each fingertip.
  • �� Aggregate tactile features through a finger- and pose-aware tactile compressor.
  • �� Inject tactile latent into a video diffusion world model.
  • �� Test on six contact-rich dexterous manipulation tasks.

Experiments

The experiments were conducted on a 22-DoF bimanual platform across six contact-rich dexterous manipulation tasks. Approximately 100 demonstrations per task were used for vision-to-touch learning, ensuring visual prediction quality within 0.5 dB. Results showed DexTacWAM achieved the highest score on every task, significantly outperforming baselines.

Results

DexTacWAM achieved the highest score on all six tasks, averaging 70.6 compared to the strongest baseline of 38.0. Removing tactile world modeling reduced the four-task mean score from 74.7 to 26.6, highlighting the importance of modeling contact evolution.

Applications

DexTacWAM is applicable to dexterous manipulation tasks requiring fine contact control, such as object grasping, handling, and assembly. It performs well under limited data and computational resources, making it suitable for industrial and service robotics.

Limitations & Outlook

DexTacWAM's performance is limited by the scarcity of tactile data, as tactile data collection is costly. The model's robustness in highly dynamic environments needs further validation. Future research could explore more efficient tactile data collection methods.

Plain Language Accessible to non-experts

Imagine you're in a kitchen, needing to use your fingers to skillfully grab and manipulate various utensils and ingredients. DexTacWAM is like a super flexible kitchen assistant that can sense the state of each object through touch, like whether it's slipping or securely held. Traditional visual assistants can only see the surface of objects, but DexTacWAM can feel the changes in contact between your fingers and the objects, just like you feel the texture of ingredients with your fingers. This way, DexTacWAM helps robots perform better in complex manipulation tasks.

ELI14 Explained like you're 14

Hey there! Have you ever wondered how robots can be as flexible as you when using their fingers to grab things? DexTacWAM is a super cool robot assistant that can not only see with its eyes but also 'feel' things with its fingers. Imagine you're playing with blocks, and DexTacWAM is like your best buddy, helping you build a taller and more stable tower. It can feel the weight and shape of each block, just like you do with your fingers. This way, it can help you build more complex block towers!

Glossary

DexTacWAM

A model that combines visual and tactile information for dexterous manipulation.

Used in the paper to model contact dynamics.

Tactile Compressor

A module that aggregates tactile information, preserving finger pose and contact information.

Used to reduce computational cost.

Video Diffusion Model

A model used to predict future visual states.

Used for joint modeling of visual and tactile dynamics.

World-Action Model (WAM)

A model combining world modeling and action generation.

Used in robot learning.

Vision-to-Touch Learning

The process of learning tactile predictive capabilities from visual data.

Used to extend pretrained video models.

Open Questions Unanswered questions from this research

  • 1 How to improve DexTacWAM's robustness in highly dynamic environments? The current model is limited in such scenarios and needs further research.
  • 2 How to reduce the cost of tactile data collection? The scarcity of tactile data limits the model's widespread application.

Applications

Immediate Applications

Industrial Robots

DexTacWAM can be used in industrial robots for precise manipulation tasks like assembly and handling, improving efficiency and accuracy.

Long-term Vision

Service Robots

In the future, DexTacWAM could be applied to service robots, helping them perform flexible operations in complex environments, like home assistants.

Abstract

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.

cs.RO cs.AI cs.CV