CNN-Based Camera Pose Estimation and Localisation of Scan Images for Aircraft Visual Inspection
CNN-based camera pose estimation for aircraft inspection achieves <0.24m and 2° error.
Xueyan Oh, Leonard Loh, Shaohui Foong et al.
CNN-based camera pose estimation for aircraft inspection achieves <0.24m and 2° error.
Xueyan Oh, Leonard Loh, Shaohui Foong et al.
Splatblox uses Gaussian Splatting for outdoor robot navigation, achieving a 50% success rate increase.
Samarth Chopra, Jing Liang, Gershom Seneviratne et al.
MobileVLA-R1 combines CoT reasoning with GRPO RL, achieving 5% improvement in quadruped robot vision-language control.
Ting Huang, Dongjian Li, Rui Yang et al.
Proposes AINA, a framework using smart glasses to learn multi-finger robot manipulation from in-the-wild human videos without robot data.
Irmak Guzey, Haozhi Qi, Julen Urain et al.
MiMo-Embodied is a unified cross-embodied foundation model integrating autonomous driving and embodied AI, achieving SOTA on 17 embodied AI and 12 autonomous driving benchmarks through multi-stage training.
Xiaoshuai Hao, Lei Zhou, Zhijian Huang et al.
Proposes LongComp with scene factorization to enhance zero-shot generalization in autonomous driving trajectory prediction, reducing OOD gap to 2.8%.
Benjamin Stoler, Jonathan Francis, Jean Oh
ViPRA leverages video prediction and latent action learning to enable robot high-frequency smooth control, achieving 16% performance gain.
Sandeep Routray, Hengkai Pan, Unnat Jain et al.
EquiTac leverages SO(2) symmetry to enhance sample efficiency in tactile policy learning, significantly reducing training samples.
Yizhe Zhu, Zhang Ye, Boce Hu et al.
X-Diffusion employs diffusion models to learn cross-embodiment robot policies from human demonstrations, achieving 16% success rate improvement.
Maximus A. Pace, Prithwish Dan, Chuanruo Ning et al.
TWIST2 integrates PICO4U VR and a custom 2-DoF neck to enable scalable, mocap-free humanoid teleoperation and data collection, achieving near 100% success in 100 demonstrations within 15 minutes.
Yanjie Ze, Siheng Zhao, Weizhuo Wang et al.
Proposes a particle-based cross-embodiment world model for dexterous manipulation, enhancing generalization to unseen hands.
Zihao He, Bo Ai, Tongzhou Mu et al.
Proposed a data-efficient flow-based equivariant grasp synthesis architecture handling diverse gripper types; dataset includes 25,000 scenes and 20 million grasps.
Roman Freiberg, Alexander Qualmann, Ngo Anh Vien et al.
Proposed a multi-modal neuro-symbolic framework combining panoramic images and 3D point clouds for spatial reasoning in robotics.
Simindokht Jahangard, Mehrzad Mohammadi, Abhinav Dhall et al.
PHUMA combines physics-aware filtering and constrained retargeting to build a 73-hour humanoid motion dataset with high physical reliability.
Kyungmin Lee, Sibeen Kim, Youngdo Lee et al.
Proposes a deep active inference framework combining diffusion policy and multi-timescale world model, improving robotic exploration and navigation.
Riko Yokozawa, Kentaro Fujii, Yuta Nomura et al.
Cosine-warming Temperature Sampling raises Libero UniVLA success from 0.76 to 0.85 under severe task imbalance.
Basavasagar Patil, Sydney Belt, Jayjun Lee et al.
ISE combines online embedding refinement and failed-plan rejection to improve video replanning on five Meta-World tasks with 400 trials each.
Po-Chen Ko, Jiayuan Mao, Yu-Hsiang Fu et al.
First efficiency-focused VLA survey: four-way taxonomy, from 55B to 0.24B models.
Weifan Guan, Qinghao Hu, Aosheng Li et al.
LIBERO-Plus systematically analyzes VLA model robustness under seven perturbations, revealing performance drops from 95% to below 30%.
Senyu Fei, Siyin Wang, Junhao Shi et al.
X-VLA employs soft-prompted Transformer, integrating diverse robotic datasets for scalable cross-embodiment learning with minimal parameter tuning.
Jinliang Zheng, Jianxiong Li, Zhihao Wang et al.