GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.RO 2609.05401

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

ROBORMBENCH reveals reward instability in vision-language models under semantically equivalent instructions, featuring 2,390 trajectories and 21,673 paraphrases.

Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al.

2026-09-05 96
cs.RO 2609.05376

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

Using ACT, this study examines visual distractors' impact on imitation policies, significantly improving UR3e robustness.

Vivek Chavan, Pengtao Xie, Yahuan Shi et al.

2026-09-05 91
cs.RO 2609.05369

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

Proposes a neuro-symbolic framework combining VLA control and task graphs for long-horizon vision-language-action manipulation.

Vivek Chavan, Yahuan Shi, Oliver Heimann et al.

2026-09-05 97
cs.RO 2609.05361

Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

Developed a humanoid robot prototype for multimodal HRI, achieving 96% gesture recognition accuracy.

Thang Tran Viet, Thanh Nguyen Canh, Huy Uong Gia et al.

2026-09-05 94
cs.RO 2609.05331

Adaptation Needs in Robotic Systems: Assessing Behavior Trees and Their Enhancement

Behavior Trees lack adaptability in robotic systems, needing enhancements for dynamic environments.

Mehran Rostamnia, Gianluca Filippone, Ricardo Caldas et al.

2026-09-05 88
cs.RO 2609.05325

FIRE-LIVWO: Robust LiDAR-Inertial-Visual-Wheel Odometry via Failure-Immune mmWave Radar Enhancement

FIRE-LIVWO achieves robust LiDAR-Inertial-Visual-Wheel Odometry with mmWave radar enhancement, average localization error of 5.677m.

Kun Hu, Menggang Li, Kaidi Wu et al.

2026-09-05 90
cs.RO 2609.05324

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA evaluates VLA models with 527K trajectories in complex scenes and long-horizon tasks.

Zhenxuan Fan, Bo Zhang, Yutong Lin et al.

2026-09-05 74
cs.RO 2609.05300

Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds

H2INT uses a two-stage Transformer to enhance robot navigation safety and robustness in dense crowds.

Ao Shen, Kaixi Chen, Shiwei Liu et al.

2026-09-04 28
cs.RO 2609.05282

Temporal Tactile Encoding and Compliance for Intent-Aware Robot-to-Human Bimanual Handover

Combining VLA model and compliance controller enhances robot-to-human bimanual handover efficiency.

Pasquale Marra, Stefano Berti, Gabriele Mario Caddeo et al.

2026-09-04 40
cs.RO 2609.03984

MulDP: Multimodal Diffusion Policy for Autonomous Quadruped Parkour Navigation across Complex Terrains

MulDP employs multimodal diffusion models to enable autonomous quadruped navigation across complex terrains, achieving 89.7% success in diverse scenarios.

Kangmai Hu, Yueqi Zhang, Peng Zhai et al.

2026-09-03 86
cs.RO 2609.03970

Automated Weld Seam Recognition and 3D Mapping for Robotic Post Processing Using Photogrammetry and Semantic Segmentation

Combining semantic segmentation and photogrammetry, the method localizes weld seams with ~4.2mm RMSE using smartphone images.

Augustin Raju, Abilash Madavath, Chandra Yuvesh Aubeeluck et al.

2026-09-03 80
cs.RO 2609.03276

R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

R2S-Eval combines real-to-sim calibration with VLM preference evaluation for stable robot behavior ranking.

Yidi Wang, Feixiang Ruan, Ruoqu Chen et al.

2026-09-03 22
cs.RO 2609.02861

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

Proposes TRACE framework with 98.6% evidence traceability, enabling end-to-end decision auditability for autonomous robots.

Cagri Temel

2026-09-03 67
cs.RO 2609.01579

SG-AMP: Scene-Graph-Guided Active Perception and Semantics-Aware Motion Planning for Pepper Plants

SG-AMP combines robust depth completion, scene graph reasoning, and semantics-aware active view planning, achieving 55.27% semantic mIoU and 38.67% PQ on pepper data.

Rohit Menon, Shiva Rudra Lolla, Niklas Mueller-Goldingen et al.

2026-09-02 80
cs.RO 2609.01518

A System for Fast, Resilient, and Adaptable Loco-Manipulation Behaviors on Humanoid Robots

A behavior architecture combining environment templates and behavior trees enables fast, resilient, and adaptable humanoid loco-manipulation with runtime editing.

Duncan Calvert, Luigi Penco, Dexton Anderson et al.

2026-09-02 80
cs.RO 2609.01453

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

This study compares expert and imitation policies using Transformer-based ACT in ParcelStow, revealing significant success rate decline at higher speeds, with expert maintaining better robustness.

Clinton Enwerem, John S. Baras, Calin Belta

2026-09-01 75
cs.RO 2609.01404

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

Proposes DroneCATS-Agent using multimodal LLMs for autonomous drone control, covering commanding, approaching, tracking, and multi-drone coordination, evaluated on a unified benchmark.

Jaewoo Park, Minyoung Lee, Sukmin Seo et al.

2026-09-01 72
cs.RO 2609.01351

Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs

Extends Rao-Blackwellized online POMDP with hybrid belief representation, reducing sampling variance for high-dimensional robotic planning.

Jiho Lee, Nisar Ahmed, Kyle Hollins Wray et al.

2026-09-01 70
cs.RO 2608.30935

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0 leverages pretrained VLM's spatial reasoning via unified pointing and RVQ trajectory decoding, enabling multi-task embodied navigation without task-specific heads.

Shaoan Wang, Aocheng Luo, Fei Huang et al.

2026-08-31 76
cs.RO 2608.30880

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

Zeva enables in-context causal learning, allowing robots to self-evolve without parameter updates, improving success rates over repeated attempts.

Fu Chen, Xin Ding, Bingjia Huang et al.

2026-08-31 80
Prev 1 2 3 4 5 ... 47 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home