ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations

TL;DR

ForgeVLA learns VLA models from distributed vision-action pairs using federated learning, without language annotations, significantly improving performance.

cs.CV 🔴 Advanced 2026-05-08 36 views
Yuhao Zhou Yunpeng Zhu Yang Zhou Jindi Lyu Jian Lan Zhangyuan Wang Dan Si Thomas Seidl Qing Ye Jiancheng Lyu
federated learning vision-language-action robotics unsupervised learning data privacy

Key Findings

Methodology

ForgeVLA is a federated learning framework that learns VLA models from distributed vision-action pairs without centralizing raw data or requiring manual annotations. Each client is equipped with an embodied instruction classifier that maps vision-action pairs to a predefined instruction set, recovering the missing language modality. To address vision-language feature collapse, ForgeVLA combines a client-side contrastive planning loss with a server-side adaptive aggregation strategy.

Key Results

  • ForgeVLA achieved a success rate of 55.2% on the Libero-Goal dataset, a 26.4% improvement over FedAvg.
  • On the LIBERO-Object dataset, ForgeVLA achieved a success rate of 98.6%, nearly matching the centralized upper bound.
  • For the Pass@50 metric, ForgeVLA achieved 100% across all datasets, far exceeding other federated baselines.

Significance

ForgeVLA is significant in both academia and industry, particularly under data privacy and heterogeneity challenges. It leverages existing vision-action logs for large-scale VLA training, avoiding high annotation costs and achieving high-quality data expansion without relying on synthetic generation.

Technical Contribution

ForgeVLA overcomes limitations of existing methods by introducing new contrastive planning loss and adaptive aggregation strategies, addressing feature collapse under cross-domain distribution shifts. These innovations provide new theoretical guarantees and engineering possibilities for federated VLA learning.

Novelty

ForgeVLA is the first to train VLA models using federated learning without language annotations. Its innovation lies in using an embodied instruction classifier to recover the language modality and addressing feature collapse with contrastive loss and adaptive aggregation strategies.

Limitations

  • Feature collapse may still occur in highly heterogeneous client environments, affecting model generalization.
  • Requires a certain amount of pretraining data to fine-tune the embodied instruction classifier.

Future Work

Future work could explore more complex task instruction sets and broader application scenarios, further optimizing contrastive loss and aggregation strategies to enhance model robustness and adaptability.

AI Executive Summary

ForgeVLA is an innovative federated learning framework designed to address the data annotation bottleneck in training vision-language-action (VLA) models. Existing methods rely on costly manual annotations, whereas ForgeVLA utilizes distributed vision-action pairs, recovering the language modality through an embodied instruction classifier to form complete VLA triplets.

The method addresses vision-language feature collapse through contrastive planning loss and adaptive aggregation strategies. Experimental results show that ForgeVLA significantly outperforms other methods across multiple benchmarks, particularly under data privacy and heterogeneity challenges.

This research provides new insights for large-scale VLA training, avoiding high annotation costs and achieving high-quality data expansion without relying on synthetic generation. Future work will explore more complex task instruction sets and broader application scenarios.

Deep Analysis

Background

Vision-Language-Action (VLA) models hold great promise for robotic intelligence but are limited by the high cost of annotated data. Existing methods like data augmentation and simulation face reality gaps, while generative models struggle with plausible action labels. Federated learning offers a natural framework under data privacy and heterogeneity constraints.

Core Problem

The core problem for VLA models is the lack of large-scale, high-quality annotated data. Robotic data is difficult to centralize due to privacy and heterogeneity issues. Existing federated methods assume fully annotated VLA triplets at each client, failing to effectively utilize unannotated vision-action logs.

Innovation

ForgeVLA uses a federated learning framework to learn VLA models from distributed vision-action pairs without centralizing raw data or manual annotations. Its innovations include using an embodied instruction classifier to recover the language modality and addressing feature collapse with contrastive loss and adaptive aggregation strategies.

Methodology

  • �� Each client is equipped with an embodied instruction classifier to map vision-action pairs to a predefined instruction set.
  • �� A contrastive planning loss is used to enhance task discriminability.
  • �� The server employs an adaptive aggregation strategy to preserve client update directions.
  • �� Extensive experiments validate the contribution of each component.

Experiments

Experiments were conducted on four LIBERO benchmarks, simulating non-i.i.d. heterogeneity. The InternVLA-M1 serves as the backbone network, with LoRA applied to fine-tune the VLM encoder. Baselines include FedAvg, FedProx, and others, with evaluation metrics of task success rate and Pass@K.

Results

ForgeVLA achieved a success rate of 55.2% on the LIBERO-Goal dataset, a 26.4% improvement over FedAvg. On the LIBERO-Object dataset, it achieved a success rate of 98.6%, nearly matching the centralized upper bound. For the Pass@50 metric, ForgeVLA achieved 100% across all datasets.

Applications

ForgeVLA is applicable in robotics fields requiring data privacy, such as manufacturing, warehousing, and healthcare. Its decentralized data handling makes it widely applicable in data-sensitive industries.

Limitations & Outlook

Feature collapse may still occur in highly heterogeneous client environments, affecting model generalization. Additionally, the embodied instruction classifier requires a certain amount of pretraining data for fine-tuning. Future work will explore more complex task instruction sets and broader application scenarios.

Plain Language Accessible to non-experts

Imagine you're in a large factory where workers need to complete tasks based on vision and action, but without clear language instructions. ForgeVLA acts like a smart assistant that infers task instructions from the workers' actions, helping them perform better. This way, the factory doesn't need to spend a lot of time and money labeling each task but uses existing work records to improve efficiency. This assistant can also share experiences between different factories, helping each one adapt quickly to new task demands.

ELI14 Explained like you're 14

Imagine you're playing a super cool robot game. The robots in the game need to complete various tasks but don't have clear instructions. ForgeVLA is like a super assistant in the game, guessing the robots' tasks from their actions and helping them complete challenges better. This way, you don't have to spend time telling each robot what to do, but let them learn how to complete tasks themselves. Isn't that amazing?

Glossary

Federated Learning

A distributed learning method that allows multiple clients to collaboratively train a model without sharing raw data.

Used in ForgeVLA to learn VLA models from distributed vision-action pairs.

Vision-Language-Action Model

A model integrating visual perception, language understanding, and action control for robotic intelligence.

The target model trained by ForgeVLA through federated learning.

Embodied Instruction Classifier

A classifier that maps vision-action pairs to a predefined instruction set, recovering the missing language modality.

Used on each client to generate complete VLA triplets.

Contrastive Planning Loss

A loss function that introduces discriminative margins between task representations to mitigate vision-language feature collapse.

Enhances task discriminability in ForgeVLA models.

Adaptive Aggregation Strategy

A server-side strategy that optimizes projection coefficients of client updates to preserve each client's update direction.

Addresses heterogeneity challenges in ForgeVLA.

Open Questions Unanswered questions from this research

  • 1 How to apply ForgeVLA to more complex task instruction sets to enhance its adaptability and robustness.
  • 2 How to further optimize contrastive loss and aggregation strategies without increasing computational costs.

Applications

Immediate Applications

Manufacturing

ForgeVLA can be used in manufacturing for robotic control, leveraging existing vision-action logs for training to enhance production efficiency.

Long-term Vision

Smart Cities

In smart cities, ForgeVLA can be used for traffic management and public safety, utilizing distributed sensor data for real-time decision-making.

Abstract

Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring annotated training data. Fortunately, vision-equipped robots deployed across various domains already produce abundant vision-action pairs that can be leveraged to scale up VLA training more efficiently. However, these raw data cannot be centrally aggregated due to various constraints and also exhibit severe heterogeneity. To address these challenges, in this paper, we propose ForgeVLA, a federated VLA training framework that learns VLA models from distributed vision-action pairs without centralizing raw data or requiring manual annotations. Specifically, each client in ForgeVLA is equipped with an embodied instruction classifier that maps vision-action pairs to a predefined instruction set, recovering the missing language modality and forming complete vision-language-action triplets. Beyond triplet construction, we also identify vision-language feature collapse as a critical challenge that has been largely overlooked in prior federated VLA research. To mitigate this issue, ForgeVLA combines a client-side contrastive planning loss with a server-side adaptive aggregation strategy to learn task-discriminative representations efficiently. Extensive experiments across multiple benchmarks show that ForgeVLA significantly outperforms other baselines, and ablation studies further validate the contribution of each component.

cs.CV cs.AI