GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.RO 2601.21712

CoFreeVLA: Short-Horizon Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation

CoFreeVLA reduces dual-arm self-collision rates from 0.54 to 0.23 and improves task success rates to 0.61 using a Vision-Language-Action model with risk estimation.

Yaohua Liu, Binkai Ou, Hengjun Zhang

2026-01-29 34
cs.RO 2601.21409

DSCD-Nav: Dual-Stance Cooperative Debate for Object Navigation

Dual-Stance Cooperative Debate (DSCD-Nav) enhances object navigation via multi-round evidence cross-checking, boosting success rate by 20% on HM3Dv2.

Weitao An, Qi Liu, Chenghao Xu et al.

2026-01-29 42
cs.RO 2601.18923

DeFM: Learning Foundation Representations from Depth for Robotics

DeFM employs self-supervised pretraining on 60M depth images, learning geometric and semantic features for robotic tasks with state-of-the-art results.

Manthan Patel, Jonas Frey, Mayank Mittal et al.

2026-01-27 42
cs.RO 2601.16212

Point Bridge: 3D Representations for Cross Domain Policy Learning

Point Bridge leverages unified point cloud representations for zero-shot sim-to-real transfer, achieving up to 44% performance gains.

Siddhant Haldar, Lars Johannsmeier, Lerrel Pinto et al.

2026-01-23 39
cs.RO 2601.16035

Collision-Free Humanoid Traversal in Cluttered Indoor Scenes

HumanoidPF encodes humanoid-obstacle relationships, enabling collision-free navigation in cluttered indoor scenes with success rate over 93%.

Han Xue, Sikai Liang, Zhikai Zhang et al.

2026-01-22 60
cs.RO 2601.15222

MonoRace: Winning Champion-Level Drone Racing with Robust Monocular AI

MonoRace employs monocular camera and IMU for champion-level drone racing at speeds up to 100 km/h without external tracking.

Stavrow A. Bahnam, Robin Ferede, Till M. Blaha et al.

2026-01-22 5 citations 50
cs.RO 2601.14945

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control

TIDAL decouples semantic reasoning and high-frequency control via dual-frequency architecture, achieving 9Hz updates and 2x performance improvement in dynamic tasks.

Yuteng Sun, Haoran Wang, Ruofei Bai et al.

2026-01-21 28
cs.RO 2601.07060

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

PALM employs structured affordance prediction and progress estimation with diffusion models, achieving 91.8% success in long-horizon robotic tasks.

Yuanzhe Liu, Jingyuan Zhu, Yuchen Mo et al.

2026-01-12 44
cs.RO 2601.03447

Cost-Effective Radar Sensors for Field-Based Water Level Monitoring with Sub-Centimeter Accuracy

Commercial FMCW millimeter-wave radar achieves sub-centimeter water level accuracy with minimal calibration, enabling autonomous monitoring.

Anna Zavei-Boroda, J. Toby Minear, Kyle Harlow et al.

2026-01-07 17
cs.RO 2601.02905

LOST-3DSG: Lightweight Open-Vocabulary 3D Scene Graphs with Semantic Tracking in Dynamic Environments

LOST-3DSG employs low-dimensional word embeddings for lightweight open-vocabulary 3D scene graphs, enabling efficient dynamic object tracking.

Sara Micol Ferraina, Michele Brienza, Francesco Argenziano et al.

2026-01-06 15
cs.RO 2512.22414

Emergence of Human to Robot Transfer in Vision-Language-Action Models

Multi-source diverse pretraining enables human-to-robot skill transfer via vision-language-action models, nearly doubling generalization performance.

Simar Kareer, Karl Pertsch, James Darpinian et al.

2025-12-27 40
cs.RO 2512.21243

LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation

LookPlanGraph uses VLM graph augmentation for instruction following in dynamic environments, improving task success rates.

Anatoly O. Onishchenko, Alexey K. Kovalev, Aleksandr I. Panov

2025-12-24 24
cs.RO 2512.20475

Drift-Corrected Monocular VIO and Perception-Aware Planning for Autonomous Drone Racing

Fusion of YOLO-based gate detection and Kalman filter drift correction enhances monocular VIO for drone racing, achieving speeds over 59 km/h.

Maulana Bisyir Azhari, Donghun Han, Je In You et al.

2025-12-24 32
cs.RO 2512.15692

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Mimic-video combines pretrained video models with inverse dynamics decoders, achieving 10x sample efficiency and 2x faster convergence in robotic control tasks.

Jonas Pai, Liam Achenbach, Victoriano Montesinos et al.

2025-12-18 115 citations 35
cs.RO 2512.13660

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

RoboTracer employs 3D-aware VLM with reinforcement fine-tuning, achieving 79.1% success in spatial reasoning tasks.

Enshen Zhou, Yibo Li, Jingkun An et al.

2025-12-16 27
cs.RO 2512.09343

Development and Testing for Perception Based Autonomous Landing of a Long-Range QuadPlane

A YOLO–TensorRT–Isaac ROS QuadPlane landing stack achieved 32.7 ms per frame, enabling over 30 FPS edge inference.

Ashik E Rasul, Humaira Tasnim, Ji Yu Kim et al.

2025-12-10 19
cs.RO 2512.11891

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer

AEGIS integrates control barrier functions into VLA models, achieving over 50% improvement in obstacle avoidance and nearly 10% higher task success.

Songqiao Hu, Zeyi Liu, Shuang Liu et al.

2025-12-10 62
cs.RO 2512.04308

ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models

ResponsibleRobotBench evaluates robot responsibility in risk environments using multimodal LLMs with 85% success rate.

Lei Zhang, Ju Dong, Kaixin Bai et al.

2025-12-04 27
cs.RO 2511.23030

DiskChunGS: Large-Scale 3D Gaussian SLAM Through Chunk-Based Memory Management

DiskChunGS achieves large-scale 3D Gaussian SLAM via chunk-based memory management, completing all KITTI sequences successfully.

Casimir Feldmann, Maximum Wilder-Smith, Vaishakh Patil et al.

2025-11-28 26
cs.RO 2511.19528

Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories

DLR framework uses information-theoretic pattern discovery to generate diverse, high-success trajectories for VLA pretraining, improving downstream generalization.

Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang et al.

2025-11-24 44
Prev 1 ... 19 20 21 22 23 24 25 ... 47 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home