GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.RO 2605.23098

UfM*: Uncertainty from Motion* for DNN Depth Estimation Using Gaussians

UfM* employs Gaussian mixture models for single-inference multiview disagreement-based depth uncertainty estimation, achieving high accuracy with minimal energy.

Soumya Sudhakar, Sertac Karaman, Vivienne Sze

2026-05-22 38
cs.RO 2605.22816

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

AwareVLN introduces self-aware reasoning for VLN, achieving NE 4.02 on R2R-CE Val-Unseen, outperforming prior SOTA.

Wenxuan Guo, Xiuwei Xu, Yichen Liu et al.

2026-05-22 254
cs.RO 2605.22812

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

GesVLA integrates gesture into Vision-Language-Action models, achieving 94.3% target grounding accuracy in complex real-world tasks.

Wenxuan Guo, Ziyuan Li, Meng Zhang et al.

2026-05-22 170
cs.RO 2605.22748

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

League-based multi-agent RL achieves 22 m/s quadrotor racing with 50% collision reduction vs. single-agent baselines.

Ismail Geles, Leonard Bauersfeld, Markus Wulfmeier et al.

2026-05-22 245
cs.RO 2605.22600

Branch-Stochastic Model Predictive Control for Motion Planning under Multi-Modal Uncertainty with Scenario Clustering

Proposed Branch-Stochastic MPC with scenario clustering improves safety and real-time performance in multi-modal uncertainty motion planning.

Zekun Xing, Ramkrishna Chaudhari, Marion Leibold et al.

2026-05-21 388
cs.RO 2605.22456

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

Steins;Gate Drive employs structured future forecasting with latency decoupling, significantly improving autonomous driving safety and responsiveness.

Anjie Qiu, Hans D. Schotten

2026-05-21 52
cs.RO 2605.22283

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

SOMA employs multi-view spatial memory to enable out-of-vision manipulation, boosting success rates by 10% in real-world tasks.

Pengteng Li, Weiyu Guo, He Zhang et al.

2026-05-21 50
cs.RO 2605.21862

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

EvoScene-VLA updates scene states via actions, boosting RoboTwin task success to 88.8%.

Chushan Zhang, Ruihan Lu, Jinguang Tong et al.

2026-05-21 26
cs.RO 2605.21258

Learning Structural Latent Points for Efficient Visual Representations in Robotic Manipulation

Proposed a structural latent point learning method, achieving a 56% task success rate on RLBench.

Yicheng Jiang, Jiaxu Wang, Junhao He et al.

2026-05-20 32
cs.RO 2605.20373

SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

SUGAR leverages human videos to learn humanoid loco-manipulation skills without task rewards or reference motions, enabling zero-shot transfer.

Tianshu Wu, Xiangqi Kong, Yue Chen et al.

2026-05-20 52
cs.RO 2605.19038

Guiding Neuro-Symbolic Scenario Generation with Spatio-Temporal Logic

Proposes STRELGen, integrating diffusion models with spatio-temporal logic for targeted safety-critical scenario synthesis.

Lorenzo Bonin, Francesco Giacomarra, Luca Bortolussi et al.

2026-05-19 41
cs.RO 2605.18074

4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving

Introduces 4DLidarOpen, a large-scale multi-modal dataset with 4D FMCW lidar velocity data, enhancing dynamic scene understanding.

Kane Qian, Xin Zhao, Yining Shi et al.

2026-05-18 40
cs.RO 2606.00054

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

This survey categorizes four main methods—latent actions, world models, 2D cues, and 3D reconstructions—for transforming human videos into robot control knowledge, enabling scalable vision-language-action learning.

Zhiyuan Feng, Qixiu Li, Huizhi Liang et al.

2026-05-18 46
cs.RO 2605.17077

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning

DeMiAn method enhances robot policy learning with dense language annotations, boosting RoboCasa success by 5 points.

Bosung Kim, Ruiyi Wang, David Acuna et al.

2026-05-17 7
cs.RO 2605.15559

NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation

NavRL++ introduces a system-level framework with perturbation-aware fine-tuning and Transformer-based temporal reasoning, achieving zero-shot sim-to-real transfer in robot navigation.

Zhefan Xu, Hanyu Jin, Kenji Shimada

2026-05-15 30
cs.RO 2605.13452

CUBic: Coordinated Unified Bimanual Perception and Control Framework

CUBic employs shared codebooks for unified bimanual perception and control, achieving 12% higher success in RoboTwin benchmarks.

Xingyu Wang, Pengxiang Ding, Jingkai Xu et al.

2026-05-13 49
cs.RO 2605.12719

A Five-Layer MLOps Architecture for Connected Automated Driving

Proposes a five-layer MLOps architecture enabling collective learning and safety assurance for autonomous vehicle fleets.

Bastian Lampe, Lutz Eckstein

2026-05-13 35
cs.RO 2605.12386

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

SafeManip uses LTLf to evaluate temporal safety in robotic manipulation, revealing task success does not equal safe execution.

Chengyue Huang, Khang Vo Huynh, Sebastian Elbaum et al.

2026-05-13 504
cs.RO 2605.12347

Real-Time Whole-Body Teleoperation of a Humanoid Robot Using IMU-Based Motion Capture with Sim2Sim and Sim2Real Validation

Real-time whole-body teleoperation using Virdyn IMU motion capture, validated on Unitree G1 robot with Sim2Sim and Sim2Real.

Hamza Ahmed Durrani, Suleman Khan

2026-05-13 226
cs.RO 2605.12167

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

MoLA uses pretrained multimodal inverse dynamics to convert imagined future videos into executable robot actions, boosting success and generalization.

Yajie Li, Bozhou Zhang, Chun Gu et al.

2026-05-12 50
Prev 1 ... 10 11 12 13 14 15 16 ... 46 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home