cs.RO 2511.16518

MiMo-Embodied: X-Embodied Foundation Model Technical Report

MiMo-Embodied is a unified cross-embodied foundation model integrating autonomous driving and embodied AI, achieving SOTA on 17 embodied AI and 12 autonomous driving benchmarks through multi-stage training.

Xiaoshuai Hao, Lei Zhou, Zhijian Huang et al.

2025-11-21 39 citations 65
cs.RO 2511.07732

ViPRA: Video Prediction for Robot Actions

ViPRA leverages video prediction and latent action learning to enable robot high-frequency smooth control, achieving 16% performance gain.

Sandeep Routray, Hengkai Pan, Unnat Jain et al.

2025-11-11 39
cs.RO 2510.27420

Towards a Multi-Embodied Grasping Agent

Proposed a data-efficient flow-based equivariant grasp synthesis architecture handling diverse gripper types; dataset includes 25,000 scenes and 20 million grasps.

Roman Freiberg, Alexander Qualmann, Ngo Anh Vien et al.

2025-10-31 1
cs.RO 2510.17315

Implicit State Estimation via Video Replanning

ISE combines online embedding refinement and failed-plan rejection to improve video replanning on five Meta-World tasks with 400 trials each.

Po-Chen Ko, Jiayuan Mao, Yu-Hsiang Fu et al.

2025-10-20 21