RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

TL;DR

RoboSPA evaluates VLA models with 527K trajectories in complex scenes and long-horizon tasks.

cs.RO 🔴 Advanced 2026-09-05 73 views
Zhenxuan Fan Bo Zhang Yutong Lin Yuqian Yuan Juekai Lin Liang Liang Zhuoyi Huang Wenqiao Zhang Juncheng Li Siliang Tang Jun Xiao Yueting Zhuang
robotics vision-language action generation dataset spatial reasoning

Key Findings

Methodology

RoboSPA is built using the SAPIEN simulator and RoboTwin 2.0 framework to evaluate VLA models in fine-grained spatial reasoning and long-horizon procedural planning. The dataset includes 10 task categories and 56 base tasks, each instantiated at five difficulty levels, yielding 280 variants. Trajectories are collected through expert-coded simulations.

Key Results

  • Result 1: On the hardest tasks, all models achieve an average success rate below 25%, with some tasks dropping to 0%.
  • Result 2: X-VLA performs best in fine-grained spatial reasoning, but its success rate drops from 41.9% at L1 to 23.9% at L5.
  • Result 3: In long-horizon procedural planning, π0.5 performs best but drops from 71.2% at L1 to 23.1% at L5.

Significance

RoboSPA provides a challenging benchmark for VLA models, promoting the development of more robust, reliable, and generalizable embodied agents. It fills the gap in existing datasets for evaluating complex spatial relations and long-horizon task execution.

Technical Contribution

RoboSPA introduces new diagnostic metrics such as Object-Normalized Target Accuracy and Progress Score, surpassing existing benchmarks' binary success rates by evaluating fine-grained spatial reasoning and long-horizon procedural planning.

Novelty

RoboSPA is the first VLA dataset to make fine-grained spatial reasoning a core evaluation dimension, offering multi-level task difficulty and detailed step-level evaluation.

Limitations

  • Limitation 1: Current VLA models perform poorly in complex spatial relations and memory-intensive tasks.
  • Limitation 2: The dataset's scene diversity may not cover all real-world complexities.

Future Work

Future work can explore enhancing VLA models' spatial reasoning capabilities and memory mechanisms, developing more robust models for complex real-world scenarios.

AI Executive Summary

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision-Language-Action (VLA) models in complex scenes and long-horizon tasks. Existing datasets mainly evaluate task completion under predefined settings, lacking insight into model reasoning under increasing spatial and procedural complexity.

RoboSPA covers 10 task categories and 56 base tasks, each instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation.

Experiments show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents.

Deep Analysis

Background

Vision-Language-Action (VLA) models integrate visual perception, language understanding, and action generation, emerging as a promising paradigm for general-purpose embodied agents. However, existing datasets and benchmarks mainly evaluate these models in structured, short-horizon tasks, lacking in-depth evaluation of complex spatial relations and long-horizon task execution.

Core Problem

Deploying VLA models in real-world scenarios requires capabilities beyond short-horizon instruction following in simple scenes. Everyday manipulation tasks pose three key challenges: fine-grained target disambiguation, temporally extended task execution, and complexity-scalable embodied reasoning.

Innovation

RoboSPA addresses the gap in existing datasets by introducing evaluation dimensions of fine-grained spatial reasoning and long-horizon procedural planning. It offers multi-level task difficulty and detailed step-level evaluation.

Methodology

  • �� Built using the SAPIEN simulator and RoboTwin 2.0 framework
  • �� Designed 10 task categories and 56 base tasks
  • �� Each task instantiated across five difficulty levels, yielding 280 variants
  • �� Collected 527K trajectories across multiple embodiments and diverse scenes
  • �� Introduced Object-Normalized Target Accuracy and Progress Score as diagnostic metrics

Experiments

Experiments were conducted on the Aloha-AgileX embodiment using clean scene data for training and evaluation. Each base task was trained with a separate model using data from all five difficulty levels and evaluated independently at each level.

Results

Experiments show that current VLA models still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. On the hardest tasks, all models achieve an average success rate below 25%.

Applications

RoboSPA provides a challenging benchmark for VLA models, promoting the development of more robust, reliable, and generalizable embodied agents. It can be used to evaluate and improve VLA models' performance in complex scenes and long-horizon tasks.

Limitations & Outlook

Current VLA models perform poorly in complex spatial relations and memory-intensive tasks. The dataset's scene diversity may not cover all real-world complexities. Future work can explore enhancing VLA models' spatial reasoning capabilities and memory mechanisms.

Plain Language Accessible to non-experts

Imagine a robot working in a kitchen. It needs to complete tasks based on visual and language instructions, like finding a specific spice jar and placing it correctly. RoboSPA is like a large kitchen simulator, testing the robot's performance in complex environments. It not only needs to find the spice jar but also remember their locations and complete tasks in a specific order. It's like a chef in a busy kitchen, handling multiple tasks simultaneously and ensuring each step is correct.

ELI14 Explained like you're 14

Imagine you're playing a super complex video game with lots of tasks, like finding hidden treasures or solving puzzles. RoboSPA is like this game, but it's designed for robots. The robot needs to complete tasks based on visual and language clues, like finding the right object among similar ones. It's like in a game where you need to find the right path based on clues. The robot also needs to remember previous steps, just like you need to remember previous clues in the game.

Glossary

Vision-Language-Action Model

A model that integrates visual perception, language understanding, and action generation for robotic tasks.

Used in the paper to evaluate robots' performance in complex tasks.

Fine-Grained Spatial Reasoning

Evaluates a model's ability to locate targets in complex spatial structures.

A core evaluation dimension of RoboSPA.

Long-Horizon Procedural Planning

Evaluates a model's ability to execute tasks with multi-step decisions.

A core evaluation dimension of RoboSPA.

Object-Normalized Target Accuracy

Measures spatial reasoning ability under varying scene complexity.

A diagnostic metric for fine-grained spatial reasoning tasks.

Progress Score

Captures partial completion in long-horizon procedural planning tasks.

A diagnostic metric for long-horizon procedural planning tasks.

Open Questions Unanswered questions from this research

  • 1 Current VLA models perform poorly in complex spatial relations and memory-intensive tasks, needing enhanced spatial reasoning capabilities and memory mechanisms.
  • 2 The dataset's scene diversity may not cover all real-world complexities, requiring broader scene coverage.

Applications

Immediate Applications

Robotic Manipulation Evaluation

RoboSPA can be used to evaluate and improve VLA models' performance in complex scenes and long-horizon tasks, aiding the development of more robust robotic systems.

Long-term Vision

General-Purpose Embodied Agents

By enhancing VLA models' spatial reasoning capabilities and memory mechanisms, RoboSPA aids in developing general-purpose embodied agents capable of autonomous operation in complex real-world scenarios.

Abstract

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, \texttt{RoboSPA} introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish \texttt{RoboSPA} as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

cs.RO cs.AI cs.CV