Vesta: A Generalist Embodied Reasoning Model
Vesta unifies localization, navigation, reasoning, and planning in a single model, outperforming state-of-the-art baselines by over 20%.
Key Findings
Methodology
Vesta builds upon Qwen3-VL-8B, integrating large-scale spatial grounding data and a simple multimodal memory system. It employs supervised fine-tuning with datasets covering detection (Objects365, COCO, LVIS), embodied interactions, and robot data, to enhance spatial, navigational, and reasoning skills. The model combines multi-view images and text, utilizing Chain-of-Thought prompting for long-horizon reasoning. Its architecture avoids multi-model stacking, reducing latency and error propagation, and emphasizes spatial grounding to improve generalization across tasks. The training emphasizes spatial and navigation tasks, with a curated corpus and real robot data, ensuring robustness in real-world scenarios.
Key Results
- Vesta outperforms single best models by over 20% and ensemble baselines by more than 10% across diverse benchmarks, demonstrating its ability to match or surpass specialist systems. In navigation, success rate reaches 54.5%, comparable to specialized models like InternVLA-N1. For long-horizon tasks such as object retrieval, counting, and memory, success rates improve by over 38%, with specific tasks showing 40.1%, 35.2%, and 38.7% increases respectively.
- In real robot experiments, Vesta increases task success by 38.3%, significantly outperforming actor-only and other planning baselines. It effectively handles memory-heavy tasks, such as finding objects in drawers, counting fruits, and remembering candy, with success rates substantially higher than baselines. These results validate the model’s ability to generalize across tasks and environments.
- Ablation studies reveal that the bias towards spatial data, simple memory design, and balanced multimodal inputs are critical for performance. The model maintains high accuracy in cross-task transfer, confirming the effectiveness of the unified training approach.
Significance
This work demonstrates that a single, unified model can effectively integrate multiple capabilities traditionally handled by separate specialized systems. It simplifies the robotic control pipeline, reduces computational costs, and minimizes error cascades. The ability to outperform specialized models in both benchmark and real-world tasks signifies a major step toward scalable, general-purpose autonomous robots. Such a model can adapt to diverse environments, perform complex reasoning, and execute long-horizon tasks, addressing longstanding challenges in robotics and AI. The approach paves the way for more versatile, efficient, and intelligent robotic systems that are easier to deploy and maintain, accelerating progress toward autonomous agents capable of operating seamlessly in open-world settings.
Technical Contribution
Vesta's key innovations include the integration of a large-scale spatial grounding dataset, a simple yet effective multimodal memory system, and a Chain-of-Thought inspired long-horizon reasoning framework. The model’s architecture avoids multi-model ensembles, reducing latency and error propagation, and employs targeted supervised fine-tuning emphasizing spatial and navigational tasks. Its ability to unify multiple capabilities within a single model, while maintaining high performance across diverse benchmarks, represents a significant advancement over prior multi-model or specialist approaches. The combination of data curation, model design, and training strategy offers a new paradigm for scalable embodied AI.
Novelty
This research is the first to successfully unify localization, navigation, embodied reasoning, and planning into a single large model trained on a curated, spatially-biased dataset. Unlike previous works that rely on multiple specialized modules, Vesta demonstrates that a single, generalist model can match or outperform task-specific systems. Its innovative use of simple multimodal memory and chain-of-thought reasoning for long-horizon tasks sets a new standard in embodied AI, bridging the gap between specialized and generalist approaches.
Limitations
- The model heavily depends on large-scale spatial grounding datasets; limited data diversity may hinder performance in unseen or highly dynamic environments.
- Inference speed remains a concern for real-time applications, especially in computationally constrained settings.
- Handling extreme out-of-distribution scenarios or novel tasks not represented in training data remains challenging, indicating room for improved adaptability.
Future Work
Future efforts will focus on optimizing inference efficiency, enabling real-time deployment. Enhancing cross-domain robustness and continual learning capabilities will be key to adapting to new environments. Expanding multimodal inputs, such as incorporating audio or tactile data, can further improve contextual understanding. Additionally, integrating online learning mechanisms will allow Vesta to adapt and improve during deployment, pushing toward truly autonomous, versatile robotic systems.
AI Executive Summary
In the quest for autonomous robots capable of operating seamlessly in complex, open-world environments, researchers have long relied on multi-model stacks combining specialized modules for localization, navigation, reasoning, and planning. While effective within narrow domains, these systems are computationally expensive, slow, and prone to cascading errors. Addressing these limitations, this work introduces Vesta, a unified embodied model that integrates all core capabilities into a single architecture. Built upon the Qwen3-VL-8B foundation, Vesta leverages a curated spatial grounding dataset, a straightforward multimodal memory system, and a chain-of-thought inspired reasoning framework to handle long-horizon tasks. The model is fine-tuned with a mixture of data emphasizing spatial, navigational, and embodied reasoning skills, ensuring broad generalization.
Extensive evaluations across diverse benchmarks demonstrate Vesta’s superiority. It surpasses individual state-of-the-art models by over 20% on average and exceeds ensemble baselines by more than 10%, confirming that a single generalist can match or outperform multiple specialists. In real-world robotic tasks, Vesta improves success rates by over 38%, particularly excelling in memory-dependent and long-horizon scenarios such as object retrieval, counting, and complex manipulation. These results highlight the potential of unified models to simplify robotic architectures, reduce computational overhead, and enhance robustness.
The significance of this research lies in its paradigm-shifting approach: moving from modular, multi-model systems toward integrated, scalable, and more reliable generalist models. This not only addresses longstanding bottlenecks but also opens new avenues for deploying autonomous agents in dynamic, unpredictable environments. Despite its promising performance, challenges remain, including dependence on large datasets, inference speed, and adaptability to unseen scenarios. Future work will focus on optimizing efficiency, expanding multimodal integration, and enabling online learning, ultimately paving the way for more intelligent, versatile, and autonomous robotic systems capable of operating in real-world complexities.
Deep Dive
Plain Language Accessible to non-experts
想象你有一个超级聪明的助手,它可以帮你做很多事情,比如找到房间里的东西、规划路线、记住你说过的话,甚至提前想好下一步要做什么。这个助手就像一个非常聪明的机器人,它用一个特别的记事本,把所有过去的事情都记下来,这样它就不会忘记。每次你告诉它一个目标,比如“帮我找苹果”,它会先看看记事本,记得你之前说过的事情,然后用它的聪明脑袋,帮你找到苹果。它还能记住很多信息,帮你规划最好的路线,甚至在你走错路时提醒你。这个助手的厉害之处在于,它不用很多不同的工具,就能完成各种任务,就像你用一个万能的工具箱,里面装满了各种工具,随时可以用。这样,机器人就变得更聪明、更可靠,可以在复杂的环境中帮你做很多事情,就像你身边的一个超级助手一样。
ELI14 Explained like you're 14
想象你有个超级朋友,他不仅能帮你做作业,还能记住你上次说的事情,比如你把书包放在哪儿了,或者你和朋友玩了什么游戏。每次你问他问题,他都能马上回答你,而且还能帮你计划下一步,比如告诉你什么时候复习、怎么安排时间。这个朋友就像Vesta一样,是个超级聪明的机器人助手。它可以同时记住很多事情,知道在哪里找到东西,怎么走,甚至能提前想好下一步要做什么。它用一种特别的“记事本”帮忙,把所有的事情都记下来,随时翻阅。这样,无论任务多复杂,它都能快速找到答案,帮你完成任务。它就像你身边的一个超级聪明的朋友,帮你解决各种难题,让生活变得更轻松、更有趣!
Glossary
空间引导数据 (Spatial Grounding Data)
用于训练模型识别环境中的物体和区域,帮助理解空间布局。
在模型空间定位和导航能力训练中使用。
链式推理 (Chain-of-Thought)
逐步推导长时依赖关系的推理策略,通过中间步骤引导最终决策。
增强模型在复杂推理任务中的表现。
多模态记忆机制 (Multimodal Memory)
结合视觉和文本信息的记忆系统,用于存储和检索历史信息。
模型中的关键组件,用于长时推理。
监督微调 (Supervised Fine-tuning)
利用标注数据对预训练模型进行针对性训练,提升特定任务性能。
模型训练的重要步骤。
长时依赖 (Long-term Dependency)
模型在推理中考虑过去长时间跨度的信息,做出合理判断。
在机器人长时任务中尤为关键。
Open Questions Unanswered questions from this research
- 1 如何提升模型在极端复杂环境中的适应性和鲁棒性,尤其在数据不足或环境剧烈变化时的表现。
- 2 模型推理速度与资源消耗的平衡,确保在实际应用中既高效又准确。
Abstract
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that consolidates these capabilities into a single foundation model. Our approach combines a diverse and massive curated corpus designed to induce spatial grounding and a simple multimodal memory harness that enables reasoning over extended time horizons. Across diverse benchmarks, Vesta on average beats individual SOTA baselines by >$20\%$ and beats an ensemble of per-category-best baselines by $>10\%$ -- thus demonstrating that a generalist model can match or exceed specialists. On real-world robotic tasks requiring memory and reasoning, Vesta improves task success by >35\%. Our work thus demonstrates that a single generalist is a feasible, scalable, and arguably preferable alternative to combining specialists.