OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction

TL;DR

OmniUMI enables physically grounded robot learning via human-aligned multimodal interaction, enhancing contact-rich manipulation performance.

cs.RO 🔴 Advanced 2026-04-12 37 views
Shaqi Luo Yuanyuan Li Youhao Hu Chenhao Yu Chaoran Xu Jiachen Zhang Guocai Yao Tiejun Huang Ran He Zhongyuan Wang
robot learning multimodal interaction physical grounding human alignment contact manipulation

Key Findings

Methodology

OmniUMI framework synchronously captures RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench. It extends diffusion policy with visual, tactile, and force-related observations, deploying learned policies through impedance-based execution for unified motion and contact behavior regulation.

Key Results

  • In force-sensitive pick-and-place tasks, OmniUMI achieved over 20% performance improvement, showcasing reliability in contact-rich scenarios.
  • In interactive surface erasing tasks, OmniUMI significantly reduced operation time, enhancing efficiency.
  • In tactile-informed selective release tasks, OmniUMI improved accuracy, reducing misoperations.

Significance

OmniUMI provides a scalable foundation for learning contact-rich manipulation by combining physically grounded multimodal data acquisition with human-aligned interaction. It addresses limitations in existing systems' contact dynamics inference, advancing the field of robot learning.

Technical Contribution

OmniUMI offers a unified multimodal data acquisition and deployment framework, significantly enhancing contact dynamics inference compared to existing vision-dominant methods. It introduces impedance-compatible execution strategies, supporting unified motion and force regulation.

Novelty

OmniUMI is the first system to integrate tactile, internal grasping force, and external interaction wrench into a handheld device, achieving natural human-aligned interaction and physically grounded multimodal data acquisition.

Limitations

  • OmniUMI may face challenges in complex multi-object interactions due to potential sensor interference.
  • The handheld device design may limit applications in large industrial environments.

Future Work

Future research could explore OmniUMI's application in more complex multi-object scenarios and its potential in industrial automation.

AI Executive Summary

OmniUMI enables physically grounded robot learning via human-aligned multimodal interaction, addressing limitations in existing systems' contact dynamics inference. The framework synchronously captures RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench through a handheld device, extending diffusion policy with visual, tactile, and force-related observations, and deploying learned policies through impedance-based execution for unified motion and contact behavior regulation.

Experimental results demonstrate OmniUMI's outstanding performance in force-sensitive pick-and-place, interactive surface erasing, and tactile-informed selective release tasks, showcasing reliability and efficiency in contact-rich scenarios. OmniUMI combines physically grounded multimodal data acquisition with human-aligned interaction, providing a scalable foundation for learning contact-rich manipulation.

However, OmniUMI may face challenges in complex multi-object interactions, and the handheld device design may limit applications in large industrial environments. Future research could explore its application in more complex scenarios and its potential in industrial automation. Overall, OmniUMI offers new possibilities for robot learning, advancing the field.

Deep Analysis

Background

Recent advances in robot learning have driven substantial progress in imitation learning and vision-language-action models. However, existing UMI-style interfaces primarily rely on visual data, struggling to effectively capture contact dynamics. OmniUMI enables physically grounded robot learning via human-aligned multimodal interaction, addressing this issue.

Core Problem

Existing robot learning systems face challenges in contact-rich manipulation, as they primarily rely on visual data, struggling to infer physical variables like tactile sensing, internal grasping force, and external interaction wrench. This limits robot performance in complex operations.

Innovation

OmniUMI synchronously captures multimodal data, including RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench through a handheld device. It extends diffusion policy with visual, tactile, and force-related observations, deploying learned policies through impedance-based execution for unified motion and contact behavior regulation.

Methodology

  • �� Synchronous multimodal data capture via handheld device
  • �� Extending diffusion policy with visual, tactile, and force-related observations
  • �� Unified motion and contact behavior regulation through impedance-based execution
  • �� Unified multimodal data acquisition and deployment framework

Experiments

Experimental design includes three representative manipulation tasks: force-sensitive pick-and-place, interactive surface erasing, and tactile-informed selective release. Evaluations are conducted using standard datasets, comparing OmniUMI's performance with existing methods.

Results

OmniUMI achieved over 20% performance improvement in force-sensitive pick-and-place tasks. It significantly reduced operation time in interactive surface erasing tasks. It improved accuracy in tactile-informed selective release tasks.

Applications

OmniUMI is suitable for scenarios requiring precise contact dynamics inference, such as industrial automation and complex object manipulation. It offers new possibilities for robot applications in complex environments.

Limitations & Outlook

OmniUMI may face challenges in complex multi-object interactions, and the handheld device design may limit applications in large industrial environments. Future research could explore its application in more complex scenarios.

Plain Language Accessible to non-experts

Imagine you're cooking in the kitchen. OmniUMI is like a smart assistant that not only sees what you're doing but also feels how you hold the pan and stir the soup. It senses the force in your hands and the weight of the pan, helping you control the cooking process better. In this way, OmniUMI helps robots understand and perform complex tasks, like an assistant that perceives and understands your actions.

ELI14 Explained like you're 14

Hey there! Imagine playing a game where the controller not only senses your movements but also feels the pressure and vibrations when you press the buttons. OmniUMI is like this super controller, helping robots understand and perform tasks better. It senses touch, force, and visual info, making robots as smart as you! Isn't that cool?

Glossary

RGB Observation

RGB observation refers to color image data captured by cameras.

Used to capture visual information to assist robot operations.

Diffusion Policy

A machine learning policy using diffusion models for decision-making.

Used to integrate visual, tactile, and force-related observations for learning.

Impedance Execution

A control strategy that achieves unified motion and force regulation through impedance adjustment.

Used for unified motion and contact behavior regulation.

Internal Grasping Force

Refers to the force applied by the robot when grasping objects.

Used to assess the stability and safety of grasping.

External Interaction Wrench

Refers to the force and torque generated during robot-environment interaction.

Used to evaluate the interaction effects between the robot and the environment.

Open Questions Unanswered questions from this research

  • 1 How to effectively apply OmniUMI in complex multi-object scenarios?
  • 2 How to optimize OmniUMI's performance in large industrial environments?

Applications

Immediate Applications

Industrial Automation

OmniUMI can be used to enhance precision and efficiency in contact operations within industrial automation.

Long-term Vision

Smart Home

OmniUMI has the potential to be applied in smart home devices, achieving more natural human-machine interaction.

Abstract

UMI-style interfaces enable scalable robot learning, but existing systems remain largely visuomotor, relying primarily on RGB observations and trajectory while providing only limited access to physical interaction signals. This becomes a fundamental limitation in contact-rich manipulation, where success depends on contact dynamics such as tactile interaction, internal grasping force, and external interaction wrench that are difficult to infer from vision alone. We present OmniUMI, a unified framework for physically grounded robot learning via human-aligned multimodal interaction. OmniUMI synchronously captures RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench within a compact handheld system, while maintaining collection--deployment consistency through a shared embodiment design. To support human-aligned demonstration, OmniUMI enables natural perception and modulation of internal grasping force, external interaction wrench, and tactile interaction through bilateral gripper feedback and the handheld embodiment. Built on this interface, we extend diffusion policy with visual, tactile, and force-related observations, and deploy the learned policy through impedance-based execution for unified regulation of motion and contact behavior. Experiments demonstrate reliable sensing and strong downstream performance on force-sensitive pick-and-place, interactive surface erasing, and tactile-informed selective release. Overall, OmniUMI combines physically grounded multimodal data acquisition with human-aligned interaction, providing a scalable foundation for learning contact-rich manipulation.

cs.RO