Robostral Navigate

TL;DR

Robostral Navigate is a 8B vision-language model using monocular RGB input, achieving 77.4% success in navigation benchmarks.

cs.RO 🔴 Advanced 2026-07-23 42 views
Abdelaziz Bounhar Abhijeet Somani Aditi Kabra Adrian Valente Adrien Petralia Adrien Sade Alan Jeffares Albert Jiang Aleksandr Timashov Alexandre Cahill Alexandre Gavaudan Alexandre Laval Alexandre Sablayrolles Amelie Heliou Amos You Andre Jonasson Andrew Bai Andrew Ehrenberg Andrew Zhao Angele Lenglemetz Anmol Agarwal Antonia Calvi Arata Suzuki Arjun Majumdar Arthur Fournier Artjom Joosen Avinash Sooriyarachchi Aylin Guliz Akkus Aysenur Karaduman Baptiste Bout Baptiste Roziere Baudouin De Monicault Benjamin Holzschuh Benjamin Lefaudeux Benjamin Tibi Bernhard Stadlbauer Blazej Osinski Camille Le Scao Chaoran Yu Charlotte Cronjager Chen-Yo Sun Chris Bamford Christian Wallenwein Christophe Renaudin Clemence Lanfranchi Corentin Barreau Corentin Sautier Cristiana-Diana Diaconu Cyprien Courtot Daniel Marczak Darius Dabert Diego de Las Casas Dominik Nuss Dylan Rubini Dzmitry Soupel Elizaveta Demyanenko Elliot Chane-Sane Emilien Fugier Emmanuel Gottlob Erik Aas Etienne Goffinet Etienne Millon Eujeong Choi Fabian Paischer Fabian Schlager Faruk Ahmed Federico Baldassarre Filip Szatkowski Florian Wiesner Gabrielle Berrada Gaetan Ecrepont Gaetan Lepage Gaspard Blanchet Gaspard Donada-Vidal Gauthier Delerce Gauthier Guinet Genevieve Hayes Georgii Novikov Giada Pistilli Gianluca Galletti Guillaume Breton Guillaume Kunsch Guillaume Lample Guillaume Martin Guillaume Raille Gunjan Dhanuka Gunshi Gupta Han Zhou Harshil Shah Hasan Furkan Vural Hedi Hadiji Hope McGovern Hugo Cisneros Hugo Thimonier Indraneel Mukherjee Ivan Cuevas Salazar Jacques Sun Jan Ludziejewski Jason Rute Jean Quentin Jean-Hadrien Chabran Jean-Malo Delignon Jie Zhang Joachim Studnia Joep Barmentlo Johannes Brandstetter John Harvill Jonas Amar Jonas Schweizer Josephine Delas Josselin Somerville Julien Denize Julien Tauran Kartik Khandelwal Khyathi Raghavi Chandu Kilian Tep Kush Jain Larissa Laich Laura Calem Laurence Aitchison Laurent Callot Laurent Fainsin Leo Cotteleer Leonard Blier Lingxiao Zhao Louis Martin Louis Serrano Lucile Saulnier Ludovic Ho Fuh Luis Montero Maarten Buyl Manon Chossegros Marcin Mozejko Margaret Jennings Markus Hennerbichler Martin Alexandre Mathieu Poiree Mathieu Schmitt Mathilde Guillaumin Matthieu Andre Matthieu Dinot Matthieu Futeral Maurits Bleeker Mauro Comi Max Mynter Maxim Berman Maxime Darrin Maxime Louis Maximilian Augustin Maximilian Muller Melina Jingting Laimon Mert Unsal Mia Chiquier Michael Pilcer Michal Pietruszka Michal Zajac Mikhail Biriuchinskii Minh-Quang Pham Minwoo Kang Morgane Riviere Namit Katariya Nathan Grinsztajn Nathan Simpson Neeraj Aggarwal Neha Gupta Ola Mysiak Oliver Leicht Olivier Bousquet Olivier Duchenne Parag Jain Patricia Wang Patrick Blies Patrick von Platen Paul Jacob Paul Wambergue Paula Kurylowicz Pavan Kumar Reddy Pavel Kuksa Philippe Pinel Philomene Chagniot Pierre Stock Pierre-Andre Savalle Piotr Milos Prateek Gupta Pravesh Agrawal Quentin Desreumaux Quentin Torroba Quercus Hernandez Ram Ramrakhya Randall Isenhour Ranjit Parva Raul Perez Pelaez Reinhard Sonnleitner Remi Delacourt Richard Kurle Rishi Shah Rob Romijnders Rohin Arora Romain Sauvestre Roman Soletskyi Rosalie Millner Rupert Menneer Sagar Vaze Samuel Barry Samuel Belkadi Samuel Humeau Sanchit Gandhi Sandeep Subramanian Sarthak Mittal Saskia Adaime Sean Cha Sebastian Kaltenbach Shashwat Dalal Shashwat Verma Sherif Waly Shrimai Prabhumoye Siddhant Waghjale Siddharth Gandhi Simon Lepage Simon Sorg Soham Ghosh Sophie Marbach Srijan Mishra Stanislas Lange Steve Hong Sumukh Aithal Szymon Antoniak Tarun Kumar Vangani Teven Le Scao Theo Cachet Thibaut Lavril Thomas Chabal Thomas Coste Thomas Defard Thomas Foubert Thomas Robert Thomas Wang Tianyu Zhang Tim Lawson Timothee Lacroix Tobias Kronlachner Tom Bewley Tom Edwards Tomas Hodan Tuhin Das Tyler Wang Ulrick BLE Umar Jamil Umberto Tomasini Valentin Mace Van Phung Vedant Nanda Victor Jouault Victor Letzelter Victor Paltz Victor Poucheret Vincent Maladiere Vincent Pfister Virgile Richard Vladislav Bataev Wassim Bouaziz Wen Ding Li William Havard William Marshall Xinghui Li Xingran Guo Xinyu Yang Yann Dreze Yassine El Ouahidi Yassir Bendou Yihan Wang Yimu Pan Yves Martin des Taillades Zaccharie Ramzi Zhenlin Xu Zsofia Csakany
robot navigation vision-language model deep learning simulation training cross-platform generalization

Key Findings

Methodology

This approach employs an 8B parameter vision-language model (VLM) combined with a pointing mechanism and a displacement fallback strategy. It leverages large-scale simulation data, generating 2.4 million trajectories across 350k scenes, trained with a prefix-caching technique that reduces token usage by 22×. The model predicts waypoints solely from monocular RGB images by inferring image coordinates and orientation, avoiding reliance on robot-specific geometry. Post-supervised training, reinforcement learning via CISPO enhances exploration and recovery. The system demonstrates strong cross-embodiment generalization, deploying on wheeled and legged robots without recalibration, setting new state-of-the-art on R2R-CE and RxR-CE benchmarks.

Key Results

  • On R2R-CE unseen validation, success rate reaches 77.4%, outperforming the previous monocular best (66.9%) by 10.5 points and multi-sensor systems (72.1%) by 5.3 points, despite using only a single RGB camera.
  • On RxR-CE, success rate is 75.1%, with an SPL of 68.7%, and a navigation error of 3.47 meters, demonstrating robust generalization across diverse language instructions and environments.
  • Training efficiency is greatly improved via prefix-caching, reducing training time from months to days, while reinforcement learning further boosts exploration and recovery, leading to more robust navigation policies.

Significance

This work addresses the longstanding challenge of reducing sensor dependency in embodied navigation, demonstrating that high-performance, scalable models can operate with minimal hardware. It paves the way for deploying advanced navigation systems on a wide range of robots, lowering costs and increasing adaptability. The combination of large-scale simulation, efficient training, and reinforcement learning offers a practical pathway toward real-world autonomous robots capable of understanding complex instructions in diverse environments.

Technical Contribution

The paper introduces a scalable navigation framework based on a large 8B vision-language model, integrating point-based waypoint prediction with a displacement fallback. It innovates with prefix-caching to drastically reduce training tokens, employs tree-based attention masks to prevent over-reliance on past actions, and combines supervised learning with online reinforcement learning for exploration. Cross-embodiment generalization is achieved through randomized robot configurations during training, enabling deployment across different robot morphologies without retraining.

Novelty

This is the first work demonstrating that a single RGB camera can achieve state-of-the-art navigation performance comparable or superior to multi-sensor systems. The key innovations include point-based waypoint inference directly in image space, a prefix-caching training method that accelerates learning, and a tree-structured attention mechanism that enhances generalization. These advances collectively redefine the minimal sensing paradigm in embodied navigation.

Limitations

  • The model's robustness diminishes under extreme lighting or highly dynamic scenes, as simulated data may not fully capture real-world variability.
  • Despite training efficiency gains, high-performance models still demand substantial GPU resources, limiting edge deployment.
  • Long-term planning and multi-modal integration remain areas for future improvement, especially in highly complex or outdoor environments.

Future Work

Future directions include integrating additional sensory modalities such as audio or tactile feedback, improving robustness in dynamic scenes, and developing online learning capabilities for continual adaptation. Exploring more efficient architectures and deployment strategies will further facilitate real-world applications, moving toward truly autonomous, general-purpose robots.

AI Executive Summary

Robotic navigation has traditionally relied on expensive sensors like depth cameras, LiDAR, or multiple cameras, which limit scalability and increase costs. These systems often require environment-specific calibration and are constrained by hardware complexity. To overcome these limitations, this research introduces Robostral Navigate, a novel approach that leverages a large-scale 8B vision-language model trained exclusively on monocular RGB images. By focusing on image-space waypoint prediction through pointing mechanisms, the system achieves robust navigation performance across diverse environments and robot types, including wheeled and legged platforms.

The core innovation lies in combining simulation-generated data, a prefix-caching training strategy, and a tree-structured attention mask. This combination drastically reduces training time—by a factor of 22—and enables the model to learn efficiently from large datasets. Post-supervised training, reinforcement learning further enhances exploration and recovery, making the system resilient to real-world challenges.

Experimental results on the R2R-CE and RxR-CE benchmarks demonstrate that Robostral Navigate surpasses existing methods, achieving success rates of 77.4% and 75.1%, respectively. Notably, it outperforms multi-sensor systems despite relying solely on a single RGB camera, highlighting the potential of minimal sensing for scalable embodied navigation.

This work signifies a paradigm shift toward cost-effective, adaptable, and scalable robotic systems. It opens avenues for deploying autonomous agents in real-world scenarios where hardware simplicity and robustness are paramount. Future research will focus on integrating additional sensory inputs, improving robustness in dynamic environments, and enabling continual learning to further advance autonomous navigation capabilities.

Deep Analysis

Background

机器人自主导航经历了从传统基于地图和激光SLAM的路径规划,到深度学习驱动的视觉理解和自然语言指令的结合。早期方法如FastSLAM和ORB-SLAM实现了高精度定位,但对环境标定和硬件成本要求较高。近年来,深度神经网络推动的VLN(Vision-and-Language Navigation)技术逐步兴起,代表作包括Zhang等的深度感知导航和Xue等的多模态融合模型。这些方法在特定环境表现良好,但大多依赖多传感器或环境预建信息,限制了其在成本敏感场景的应用。随着机器人硬件多样化,单目视觉因成本低、普及率高,成为研究重点。如何在只用单目RGB的情况下实现高效、鲁棒的导航,成为当前研究的热点。

Core Problem

现有视觉导航系统多依赖深度信息、多摄像头或预建地图,导致硬件成本高、环境适应性差。单目视觉虽成本低,但深度估计不准确、尺度变化大,难以实现鲁棒导航。如何在只用单一RGB图像的情况下,保持高成功率和路径效率,成为核心难题。此外,训练数据获取成本高,模型泛化能力不足,也限制了实际应用。解决这些问题,需突破传统多传感器依赖,设计通用、高效、易部署的导航策略。

Innovation

本研究的创新点包括:1)提出仅用单目RGB图像的端到端导航模型,避免依赖深度和多摄像头,极大降低硬件成本;2)引入点指示机制,将目标位置转化为图像空间的点,增强模型跨场景和平台的泛化能力;3)采用prefix-caching训练策略,显著缩短训练时间;4)结合树状注意力掩码,限制模型只关注必要的前序信息,避免过拟合历史动作;5)利用模拟大规模数据训练,结合强化学习优化探索和恢复能力,整体性能优于多传感器系统。

Methodology

  • �� 输入:自然语言指令与连续单目RGB图像序列。
  • �� 视觉编码:用视觉编码器提取图像特征,转化为视觉标记。
  • �� 指令融合:将指令与视觉标记拼接,作为VLM输入。
  • �� 目标点预测:模型在当前视野中推断目标点的图像坐标(u,v)和朝向变化(Δθ),避免依赖几何参数。
  • �� 偏移回退:目标不在视野时,模型预测相对平移(Δx, Δy)和旋转(Δθ)以确保连续导航。
  • �� 训练优化:采用prefix-caching,将完整轨迹打包成单一序列,减少重复编码。
  • �� 树状注意力:限制模型只关注必要的前序信息,增强泛化。
  • �� 强化学习:在模拟环境中在线优化探索和恢复能力,提升成功率。

Experiments

模型在模拟环境中生成2.4百万轨迹,涵盖350k场景,确保多样性和泛化能力。评估在R2R-CE和RxR-CE两个基准上,比较单目RGB与多传感器系统的性能差异。关键指标包括成功率(SR)、路径效率(SPL)和导航误差。超参数方面,采用Adam优化器,学习率调节,训练时间由数月缩短到数日。对比不同训练策略验证效率和性能提升。

Results

在R2R-CE验证集,成功率77.4%,超越最优单目模型(66.9%)10.5个百分点,优于多传感器系统(72.1%)5.3个百分点。RxR-CE中,成功率75.1%,路径误差仅3.47米。训练时间由数月缩短至数日,强化学习显著提升探索与恢复能力,模型在复杂环境中表现出强大鲁棒性。

Applications

该模型适用于仓储、配送、巡检等场景,成本低、部署快。只需单目RGB摄像头,无需环境标定,能自主在未知环境中导航,提升效率与安全。未来结合多模态信息,将拓展到动态环境和多任务场景,推动工业和服务机器人普及。

Limitations & Outlook

模型在极端光照或动态场景表现仍有限,模拟训练数据多样性不足可能影响实际应用。高性能模型对GPU资源需求大,限制边缘设备部署。长远规划和多模态融合仍需优化,未来需增强环境适应性和自主学习能力。

Plain Language Accessible to non-experts

想象你在一个大厨房里做菜,所有厨具和食材都在不同的地方。你只用一只眼睛(单目相机)观察厨房,看到的只是一些菜品和厨具的图片,没有深度信息。你需要根据指示找到某个食材,比如“找红色的番茄”,但只看图片很难知道它在多远。于是,你学会了用图片中的位置点指示目标,遇到看不见的目标时,就用转身或走动的方式去寻找。这个过程就像用眼睛在厨房里“点菜”,不依赖特殊传感器,只用普通的相机,就能找到目标。通过模拟大量厨房场景训练,你变得越来越聪明,能在不同厨房里自如操作,不用每次都重新调试设备。这就像你用一只眼睛在厨房里找到所有食材一样简单。

ELI14 Explained like you're 14

想象你在玩一个超级复杂的寻宝游戏,但你只有一只眼睛看着房间里的东西,没有用到特殊的雷达或多角度摄像头。你需要找到一个藏起来的玩具,指示会告诉你它在房间的哪个角落,但你看不到整个房间,只能看到一部分。于是,你学会了用手指指向图片上的目标位置,或者转身去看其他角落。每次你成功找到目标,都会记住这个位置,下次就能更快找到。这个游戏就像机器人用一只普通相机在房间里找东西,不依赖复杂的设备,只靠聪明的学习和指示,就能完成任务。通过大量模拟训练,你变得像个专业的寻宝高手,不管房间多大、多复杂,都能找到目标。

Abstract

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.

cs.RO cs.AI