FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation
FlowCorrect uses VR interface for human corrections, boosting robotic manipulation success to 80%.
Edgar Welte, Yitian Shi, Rosa Wolf et al.
FlowCorrect uses VR interface for human corrections, boosting robotic manipulation success to 80%.
Edgar Welte, Yitian Shi, Rosa Wolf et al.
Transformer-based foundation models enable full-stack transfer in robotics, covering language, vision, and motor skills.
Freek Stulp, Samuel Bustamante, João Silvério et al.
LessMimic leverages Distance Field (DF) for long-horizon humanoid interaction, enabling reference-free inference and skill composition with 80-100% success across object scales.
Yutang Lin, Jieming Cui, Yixuan Li et al.
Proposes a smooth, efficient, vectorizable contact manifold generation framework combining analytical SDF primitives and novel edge-edge collision routines.
Onur Beker, Andreas René Geist, Anselm Paulus et al.
LoTIS employs image-space localization of reference trajectories using Transformer architecture, achieving 94-98% success in diverse environments without robot-specific training.
Finn Lukas Busch, Matti Vahs, Quantao Yang et al.
SimVLA simplifies Vision-Language-Action models, achieving 98.6% success on LIBERO benchmarks with only 0.5B parameters, outperforming multi-billion models.
Yuankai Luo, Woping Chen, Tong Liang et al.
Proposed a low-cost sim-to-real transfer method by learning force direction, significantly improving contact-rich task success rates.
Yifei Yang, Anzhe Chen, Zhenjie Zhu et al.
INHerit-SG introduces hierarchical semantic scene graphs with RAG-style retrieval, significantly improving complex query handling in robot navigation.
YukTungSamuel Fang, Zhikang Shi, Jiabin Qiu et al.
ALOE integrates Q-chunking and conservative value aggregation for off-policy evaluation, boosting real-world VLA policy stability and performance.
Rushuai Yang, Hecheng Wang, Zhichao Wu et al.
LDA-1B employs a unified multi-task framework with structured DINO latent space and multimodal diffusion transformer, scaling to 1.6B parameters with over 30k hours of heterogeneous embodied data, outperforming prior methods by up to 21%.
Jiangran Lyu, Kai Liu, Xuheng Zhang et al.
ABot-N0 employs a hierarchical ‘Brain-Action’ architecture, unifying 5 navigation tasks with 16.9M trajectories, achieving SOTA performance.
Zedong Chu, Shichao Xie, Xiaolong Wu et al.
LAP enables zero-shot cross-embodiment robot control by representing low-level actions as natural language, achieving over 50% success rate.
Lihan Zha, Asher J. Hancock, Mingtong Zhang et al.
UniVTAC creates a simulation platform supporting three tactile sensors, training a multi-task encoder that improves tactile perception by 17.1% on benchmarks and 25% in real robots.
Baijun Chen, Weijie Wan, Tianxing Chen et al.
MINT uses multi-scale spectral decomposition to disentangle behavior intent from execution details, improving transfer and generalization in imitation learning.
Renming Huang, Chendong Zeng, Wenjing Tang et al.
BridgeV2W integrates embodiment masks into pretrained video models via ControlNet, improving multi-view robustness and cross-robot generalization.
Yixiang Chen, Peiyan Li, Jiabing Yang et al.
GRU-based control of a dual-segment continuum robot achieved position/orientation RMSEs of 1.11mm/4.62°, outperforming LSTM and others.
Yuancheng Shao, Yao Zhang, Jia Gu et al.
TIC-VLA introduces a latency-aware semantic control interface, enabling robust robot navigation in dynamic environments despite multi-second reasoning delays.
Zhiyu Huang, Yun Zhang, Johnson Liu et al.
Fly0 introduces a structured semantic-geometric interface for zero-shot UAV navigation, achieving over 20% success rate improvement and halving localization error.
Zhenxing Xu, Yihong Lu, Weidong Bao et al.
ConLA employs contrastive disentanglement to learn pure latent actions from human videos, surpassing robot trajectory pretraining with only video data.
Weisheng Dai, Kai Lan, Jianyi Zhou et al.
Proposes SIDP, integrating reward-guided self-imitation with diffusion models, achieving 2.5× faster inference and improved robustness in visual navigation.
Runhua Zhang, Junyi Hou, Changxu Cheng et al.