GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

TL;DR

GALA enhances vision-language-action model pretraining with geometry-aware latent action modeling, achieving 68.3% success on RoboCasa-GR1.

cs.RO 🔴 Advanced 2026-09-19 12 views
Yichen Liu Puzhen Yuan Xiang Zhu Yanjiang Guo Jianyu Chen
latent action model vision-language-action geometry-aware cross-embodiment pretraining

Key Findings

Methodology

GALA is a geometry-aware latent action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. It introduces the Unified End-effector Motion Representation (UEMR), combining visual and geometric latent actions for effective supervision across multi-embodiment data.

Key Results

  • GALA achieved a 68.3% success rate on RoboCasa-GR1 and a 75.5% real-world success rate, demonstrating its effectiveness in modeling generalizable fine-grained motions across embodiments.
  • It excelled in fine-grained motion probing and cross-embodiment retrieval, significantly outperforming existing methods.
  • Ablation studies confirmed the performance enhancement due to UEMR design.

Significance

GALA addresses the limitations of existing image-based latent action models in capturing fine-grained end-effector motions, particularly at the finger level. By introducing geometry-aware modeling, it enhances cross-embodiment pretraining generalizability and provides a new approach for joint learning from human and robot data.

Technical Contribution

GALA's technical contribution lies in its innovative combination of visual and geometric latent actions, providing a unified end-effector motion representation. Unlike existing methods, GALA captures fine-grained geometric changes without relying on embodiment-specific alignment.

Novelty

GALA is the first to incorporate 3D geometric motion into latent action modeling, addressing the inability of image-based models to capture fine-grained end-effector motions. Its novelty lies in the introduction of UEMR, enhancing cross-embodiment generalizability.

Limitations

  • GALA's computational cost is high when processing complex multi-embodiment data, potentially affecting real-time performance.
  • The model may underperform on certain specific embodiments, requiring further optimization.

Future Work

Future work could explore reducing computational costs to improve real-time performance and adaptability. Additionally, extending GALA to more diverse robotic platforms and task scenarios could be beneficial.

AI Executive Summary

GALA is a geometry-aware latent action modeling framework designed to tackle the challenges posed by heterogeneous action spaces in multi-embodiment data. Existing image-based latent action models struggle to capture fine-grained end-effector motions, especially at the finger level. GALA enhances latent action representation by introducing 3D end-effector geometric motion.

The core innovation of GALA is the introduction of the Unified End-effector Motion Representation (UEMR), which retains fine-grained motion information while improving cross-embodiment generalizability. By combining visual and geometric latent actions, GALA enables effective vision-language-action pretraining across multi-embodiment data.

Experimental results show that GALA excels in fine-grained motion probing, cross-embodiment retrieval, and downstream vision-language-action evaluation. It achieved a 68.3% success rate on RoboCasa-GR1 and a 75.5% real-world success rate, demonstrating its effectiveness in modeling generalizable fine-grained motions across embodiments. Future work will explore reducing computational costs to improve real-time performance and adaptability.

Deep Analysis

Background

The challenge of heterogeneous action spaces in multi-embodiment data complicates the learning of large-scale vision-language-action models. Existing latent action models can learn embodiment-agnostic action representations from diverse video data but struggle to capture fine-grained end-effector motions, especially at the finger level in human and dexterous robot hands.

Core Problem

The core problem is how to effectively pretrain vision-language-action models on multi-embodiment data. Due to heterogeneous action spaces, existing latent action models fail to capture fine-grained end-effector motions, particularly at the finger level.

Innovation

GALA enhances latent action representation by introducing 3D end-effector geometric motion. Its core innovation is the Unified End-effector Motion Representation (UEMR), which retains fine-grained motion information while improving cross-embodiment generalizability.

Methodology

  • �� GALA combines visual and geometric latent actions for effective supervision across multi-embodiment data.
  • �� Introduces the Unified End-effector Motion Representation (UEMR) to retain fine-grained motion information.
  • �� Enhances image-based latent actions with 3D end-effector geometric motion.

Experiments

The experimental design includes fine-grained motion probing, cross-embodiment retrieval, and downstream vision-language-action evaluation. Tests were conducted on RoboCasa-GR1 and real-world environments to validate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments.

Results

GALA achieved a 68.3% success rate on RoboCasa-GR1 and a 75.5% real-world success rate. The results demonstrate GALA's excellence in fine-grained motion probing and cross-embodiment retrieval, significantly outperforming existing methods.

Applications

GALA can be applied in vision-language-action model pretraining for multi-embodiment robotic systems, particularly in scenarios requiring fine-grained end-effector motion capture.

Limitations & Outlook

GALA's computational cost is high when processing complex multi-embodiment data, potentially affecting real-time performance. Additionally, the model may underperform on certain specific embodiments, requiring further optimization.

Plain Language Accessible to non-experts

Imagine you're in a kitchen cooking multiple dishes at once. GALA is like a smart assistant that not only helps you monitor each pot but also adjusts the heat based on the food's condition. By observing changes in the pots, it learns how to cook delicious meals on different cookware. Even if the cookware varies, it finds common cooking techniques to ensure each dish reaches its best state.

ELI14 Explained like you're 14

Imagine you're playing a super cool robot game! GALA is like your secret weapon, helping you control all sorts of robots, whether they're humanoid or mechanical hands. It understands each robot's moves and helps you make the most precise actions in the game. It's like having an unbeatable assistant in the game, helping you win every match!

Glossary

Latent Action Model

A model that infers action representations from videos, enabling learning without explicit control labels.

Used to learn embodiment-agnostic action representations from multi-embodiment data.

Unified End-effector Motion Representation

A representation that retains fine-grained motion information, enhancing cross-embodiment generalizability.

Used to enhance GALA's latent action representation capabilities.

Fine-grained Motion Probing

Evaluates a model's ability to capture subtle end-effector motions.

Used to validate GALA's effectiveness in fine-grained motion modeling.

Cross-embodiment Retrieval

Evaluates a model's ability to capture shared motion semantics across embodiments.

Used to validate GALA's generalizability in cross-embodiment action representation.

Vision-Language-Action Model

A model that combines vision, language, and action information for multimodal learning of complex tasks.

GALA is used to enhance the pretraining capabilities of this model.

Open Questions Unanswered questions from this research

  • 1 How can GALA's real-time performance and adaptability be improved without increasing computational costs?
  • 2 How can GALA be further optimized for specific embodiments where performance is suboptimal?

Applications

Immediate Applications

Robotic Systems

GALA can enhance the pretraining capabilities of vision-language-action models in multi-embodiment robotic systems, especially in scenarios requiring fine-grained end-effector motion capture.

Long-term Vision

Multimodal Learning

GALA's success could drive advancements in multimodal learning, enabling robots to better understand and execute complex tasks.

Abstract

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.

cs.RO cs.CV