GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.RO 2604.10809

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations

WARPED generates wrist-view data from monocular RGB videos, reducing data collection time by 5-8x for robot learning.

Harry Freeman, Chung Hee Kim, George Kantor

2026-04-13 26
cs.RO 2604.10647

OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction

OmniUMI enables physically grounded robot learning via human-aligned multimodal interaction, enhancing contact-rich manipulation performance.

Shaqi Luo, Yuanyuan Li, Youhao Hu et al.

2026-04-12 35
cs.RO 2604.10579

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence

AffordGen generates diverse demonstrations using 3D generative models to enhance robot manipulation generalization.

Jiawei Zhang, Kaizhe Hu, Yingqian Huang et al.

2026-04-12 37
cs.RO 2604.08726

Task-Aware Bimanual Affordance Prediction via VLM-Guided Semantic-Geometric Reasoning

Hierarchical VLM-guided semantic-geometric framework improves task-aware bimanual grasping, achieving 88.9% strategy alignment in real-world tests.

Fabian Hahne, Vignesh Prasad, Georgia Chalvatzaki et al.

2026-04-10 63
cs.RO 2604.08508

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation

Sumo method enables quadruped robots to dynamically manipulate through sample-based planning, solving multi-task problems.

John Z. Zhang, Maks Sorokin, Jan Brüdigam et al.

2026-04-10 20
cs.RO 2604.07774

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning

RoboAgent employs capability chaining within a single VLM, decomposing complex embodied tasks into basic vision-language problems, enhancing long-term planning.

Peiran Xu, Jiaqi Zheng, Yadong Mu

2026-04-09 30
cs.RO 2604.07607

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

EgoVerse unifies large-scale egocentric human demonstration data, enabling cross-embodiment robot learning with 80k episodes, 1965 tasks, and multi-institution collaboration.

Ryan Punamiya, Simar Kareer, Zeyi Liu et al.

2026-04-09 47 citations 51
cs.RO 2604.07335

TAMEn: Tactile-Aware Manipulation Engine for Closed-Loop Data Collection in Contact-Rich Tasks

TAMEn enhances bimanual manipulation success from 34% to 75% through tactile-aware closed-loop data collection.

Longyan Wu, Jieji Ren, Chenghang Jiang et al.

2026-04-09 0
cs.RO 2604.04161

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models

Proposed AAC algorithm uses action entropy to dynamically adjust chunk size, improving robotic task success rates by 2.3%.

Yuanchang Liang, Xiaobo Wang, Kai Wang et al.

2026-04-06 30
cs.RO 2604.03999

Dynamic Whole-Body Dancing with Humanoid Robots -- A Model-Based Control Approach

A model-based full-body dance generation and control framework combining MoCap, QP, TO, and MPC achieves dynamic stability and expressive motion in humanoid robots.

Shibowen Zhang, Jiayang Wu, Guannan Liu et al.

2026-04-05 51
cs.RO 2604.03497

Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving

Proposes Sim2Real-AD, a modular framework combining GOB, PAM, TPT, for zero-shot transfer of VLM-guided RL policies to real vehicles, validated on Ford E-Transit.

Zilin Huang, Zhengyang Wan, Zihao Sheng et al.

2026-04-04 47
cs.RO 2604.03181

SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

SpatialVAM achieves data-efficient robot policy learning via 3D video diffusion, improving Meta-World success rates by 22%.

Peiyan Li, Yixiang Chen, Yuan Xu et al.

2026-04-04 25
cs.RO 2604.03037

ARM: Advantage Reward Modeling for Long-Horizon Manipulation

Advantage Reward Modeling (ARM) uses tri-state labels and a MIMO Transformer to estimate relative advantage, achieving 99.4% success in long-horizon towel-folding tasks.

Yiming Mao, Zixi Yu, Weixin Mao et al.

2026-04-03 43
cs.RO 2604.01158

SMASH: Mastering Scalable Whole-Body Skills for Humanoid Ping-Pong with Egocentric Vision

SMASH integrates onboard egocentric vision and scalable whole-body skill learning, enabling humanoid robots to perform continuous outdoor ping-pong without external sensors.

Junli Ren, Yinghui Li, Kai Zhang et al.

2026-04-02 51
cs.RO 2604.01259

Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models

Proposes Bench2Drive-VL, integrating DriveCommenter for real-time closed-loop autonomous driving evaluation with multimodal reasoning.

Xiaosong Jia, Yuqian Shao, Zhenjie Yang et al.

2026-04-01 70
cs.RO 2604.00202

DreamControl-v2: Simpler and Scalable Autonomous Humanoid Skills via Trainable Guided Diffusion Priors

DreamControl-v2 trains guided diffusion models directly in robot space, integrating diverse datasets to enhance scalability and automation for humanoid skills.

Sudarshan Harithas, Sangkyung Kwak, Pushkal Katara et al.

2026-04-01 57
cs.RO 2603.29192

Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning

GenSplat enhances view generalization in robotic policy learning using 3D Gaussian Splatting.

Sen Wang, Huaiyi Dong, Jingyi Tian et al.

2026-03-31 23
cs.RO 2603.28565

StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation

StreamingVLA achieves 2.4× speedup and 6.5× halting reduction via Action Flow Matching and Adaptive Early Observation.

Yiran Shi, Dongqi Guo, Tianchen Zhao et al.

2026-03-30 27
cs.RO 2603.26666

VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation

VLA-OPD combines SFT and RL via Reverse-KL distillation, improving sample efficiency and robustness for Vision-Language-Action models.

Zhide Zhong, Haodong Yan, Junfeng Li et al.

2026-03-28 23
cs.RO 2603.26320

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching

DFM-VLA refines robot action sequences iteratively via discrete flow matching, achieving top performance on CALVIN benchmarks.

Jiayi Chen, Wenxuan Song, Jiaxin Fang et al.

2026-03-27 27
Prev 1 ... 14 15 16 17 18 19 20 ... 47 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home