ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware
ActiveScale enhances robot active perception through model, data, and hardware design, significantly improving task success rates.
Key Findings
Methodology
ActiveScale framework advances active perception through coordinated model, data, and hardware designs. The model augments VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support coherent scene understanding. Data-wise, a scalable human-robot mid-training recipe using 1000 hours of egocentric and robotic data adapts the model to temporal inputs and pose supervision. Hardware-wise, the Active-perception Mobile-manipulation Platform (AMP) supports single-operator teleoperation.
Key Results
- Experiments show ActiveScale improves success rates on active perception tasks by 30%, outperforming baseline models.
- Human-robot co-training significantly enhances model adaptability in complex environments, increasing success rates by 20%.
- Ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training, with superior performance across five tasks.
Significance
This research provides an integrated foundation for studying and developing active perception in robotic manipulation through coordinated model, data, and hardware designs. It addresses the issue of incomplete information from fixed viewpoints, enhancing task execution in dynamic, complex environments, with significant impact on academia and industry.
Technical Contribution
Proposed a pose-grounded temporal VLA for active perception and cross-view reasoning, achieving state-of-the-art performance across five real-world tasks. Introduced a scalable human-robot mid-training recipe using large-scale egocentric data, expanding active-perception training to 1000 hours of data, providing a scalable path for transferring human manipulation and active-view priors to robot policies.
Novelty
ActiveScale is the first to extend active perception through coordinated model, data, and hardware design. It uniquely enables cross-view reasoning using historical video observations and pose tokens, compared to existing methods.
Limitations
- In dynamic environments, the model may fail in complex occlusion scenarios, requiring further optimization.
- High hardware platform costs limit widespread application.
- Dataset diversity needs enhancement to improve model generalization.
Future Work
Future work can explore more dynamic environment active perception tasks, optimize model performance in complex occlusion scenarios, and reduce hardware costs to promote widespread application.
AI Executive Summary
In complex environments, robots often face incomplete information due to fixed viewpoints, affecting task success rates. Existing methods struggle to effectively acquire task-relevant information in dynamic settings. The ActiveScale framework enhances robot active perception through coordinated model, data, and hardware designs. The model uses historical video observations and explicit camera-pose supervision to support cross-view reasoning. Data-wise, a scalable human-robot mid-training recipe using 1000 hours of egocentric and robotic data adapts the model to temporal inputs and pose supervision. Hardware-wise, the Active-perception Mobile-manipulation Platform (AMP) supports single-operator teleoperation. Experimental results show ActiveScale significantly improves success rates across five real-world tasks, validating the contributions of camera-pose-aware modeling and egocentric mid-training. Future work will explore more dynamic environment active perception tasks, optimize model performance in complex occlusion scenarios, and reduce hardware costs to promote widespread application.
Deep Analysis
Background
Active perception in robotic manipulation is crucial for addressing the issue of incomplete information from fixed viewpoints. Recent advances in vision-language-action models (VLAs) have significantly broadened robots' ability to perform diverse tasks across environments. However, as tasks become more complex, observations from a fixed viewpoint are often only partially informative. Successfully completing a task may require the robot to actively acquire task-relevant information in cluttered environments, for example, by reorienting its egocentric view to inspect a bowl hidden inside a cabinet or the contents of a backpack, thereby reducing occlusion.
Core Problem
Active perception tasks require a model to infer its camera pose and predict subsequent camera motion, with the latter contingent on the former. However, existing VLAs lack these capabilities. Although egocentric data naturally contains camera motion, existing methods do not effectively exploit this signal for learning active perception. Current approaches predominantly study active perception in static environments, leaving a substantial gap to real-world mobile manipulation, where both the robot and its observation viewpoint must adapt to dynamic, cluttered settings.
Innovation
The ActiveScale framework advances active perception through coordinated model, data, and hardware designs. The model augments VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support coherent scene understanding. Data-wise, a scalable human-robot mid-training recipe using 1000 hours of egocentric and robotic data adapts the model to temporal inputs and pose supervision. Hardware-wise, the Active-perception Mobile-manipulation Platform (AMP) supports single-operator teleoperation.
Methodology
- �� Model Design: Extend input from a single image to a video sequence containing historical frames. Introduce learnable camera tokens to predict translation, rotation, and field of view. • Data and Mid-training: Mid-train on large-scale egocentric and robot data, totaling over 1000 hours. • Hardware Platform: Develop AMP integrating bimanual manipulation, independently actuated camera, and mobile-base control. A single operator can jointly control viewpoint, manipulation, and base motion, providing a scalable interface.
Experiments
Experimental design includes five real-world tasks covering container and articulated occlusion, below-table perception, multi-stage viewpoint transitions, and large vertical workspace variation. Each task uses 150 demonstrations for post-training and evaluates each method over 20 real-world rollouts under matched conditions. A rollout succeeds only if all required stages and the task-specific final-state criterion are satisfied. We report strict task success rate (SR) and task progress (TP).
Results
ActiveScale achieves higher success rates across all five task families, increasing mean SR by 30%, outperforming baseline models. Human-robot co-training significantly enhances model adaptability in complex environments, increasing success rates by 20%. Ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training, with superior performance across five tasks.
Applications
ActiveScale framework can be applied to robotic manipulation tasks in dynamic, complex environments. Suitable for scenarios requiring active acquisition of task-relevant information, such as inspecting hidden items or adjusting viewpoints to reduce occlusion. Industry impact includes improved efficiency and accuracy of automated operations.
Limitations & Outlook
The model may fail in complex occlusion scenarios, requiring further optimization. High hardware platform costs limit widespread application. Dataset diversity needs enhancement to improve model generalization.
Plain Language Accessible to non-experts
Imagine a robot working in a kitchen. It needs to find a bowl hidden inside a cabinet, but a fixed viewpoint makes it invisible. ActiveScale is like giving the robot a flexible camera that can move its viewpoint, just like a person checking different angles. By learning how humans move and operate in the kitchen, the robot can more intelligently find the bowl and complete the task. This system acts like a smart assistant, helping the robot find what it needs in complex environments.
ELI14 Explained like you're 14
Hey, friends! Imagine you're playing a game where you need to find a treasure hidden in a room. A regular robot is like a fixed camera that can only see one angle. ActiveScale is like a super camera that can move its viewpoint, just like you looking around. It learns how humans move and operate in a room, helping the robot find the treasure and win the game. Isn't that cool? This system makes robots smarter, able to find what they need in complex environments.
Glossary
Vision-Language-Action Model
A model combining vision, language, and action for robotic task execution.
Used to support robots in performing complex tasks across environments.
Egocentric Data
Data collected from a first-person perspective, often containing camera motion information.
Used to train models for active perception tasks in dynamic environments.
Active Perception
The ability of robots to actively acquire task-relevant information by adjusting viewpoints.
Addresses the issue of incomplete information from fixed viewpoints, improving task success rates.
Ablation Study
A research method that evaluates the contribution of model components by removing or altering them.
Used to validate the contributions of camera-pose-aware modeling and egocentric mid-training.
Mobile Manipulation Platform
A hardware platform supporting robotic movement and manipulation in dynamic environments.
Used to enable single-operator teleoperation for active perception tasks.
Open Questions Unanswered questions from this research
- 1 How to optimize active perception models in more complex dynamic environments? Current methods may fail in complex occlusion scenarios, requiring further research.
- 2 How to reduce hardware platform costs to promote widespread application? High costs currently limit its application in the industry.
Applications
Immediate Applications
Robotic Operations in Dynamic Environments
ActiveScale can help robots perform tasks in dynamic, complex environments, improving the efficiency and accuracy of automated operations.
Long-term Vision
Intelligent Robotic Assistant
With further optimization and cost reduction, ActiveScale could become an intelligent robotic assistant, widely used in homes and industries.
Abstract
Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a framework that advances active perception through coordinated model, data, and hardware designs. Our model augments a VLA with historical video observations and explicit camera-pose supervision, using per-frame pose tokens and a lightweight prediction head to associate observations across viewpoints and support a coherent understanding of the scene. To learn from the camera motion naturally present in human activity, we introduce a scalable human--robot mid-training recipe using 1000 hours of egocentric and robotic data, adapting the model to temporal inputs and pose supervision. We further introduce Active-perception Mobile-manipulation Platform (AMP), a robotic platform that supports active perception and mobile manipulation through single-operator teleoperation, enabling scalable collection of demonstrations that coordinate viewpoint changes and manipulation. Experiments demonstrate improved success rates on active-perception tasks, while ablation studies validate the contributions of camera-pose-aware modeling and egocentric mid-training. Together, these components provide an integrated foundation for studying and developing active perception in robotic manipulation.