Scaling Cross-Embodiment World Models for Dexterous Manipulation
Proposes a particle-based cross-embodiment world model for dexterous manipulation, enhancing generalization to unseen hands.
Zihao He, Bo Ai, Tongzhou Mu et al.
Proposes a particle-based cross-embodiment world model for dexterous manipulation, enhancing generalization to unseen hands.
Zihao He, Bo Ai, Tongzhou Mu et al.
GUI-AIMA aligns multimodal attention for efficient GUI grounding, achieving 61.5% average accuracy.
Shijie Zhou, Viet Dac Lai, Hao Tan et al.
SpecDiff-2 combines discrete diffusion models with alignment techniques, achieving up to 5.5× speed-up in LLM inference without accuracy loss.
Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen et al.
VinciCoder unifies multimodal code generation with coarse-to-fine visual reinforcement learning, enhancing code executability and visual fidelity.
Xuanle Zhao, Deyang Jiang, Zhixiong Zeng et al.
Proposes an object-aware 4D human motion generation framework based on 3D Gaussian representations and motion diffusion priors.
Shurui Gui, Deep Anil Patel, Xiner Li et al.
Phased DMD introduces phase-wise distillation with subinterval score matching, boosting multi-step generative performance, validated on image/video synthesis with large-scale models.
Xiangyu Fan, Zesong Qiu, Zhuguanyu Wu et al.
Proposed a data-efficient flow-based equivariant grasp synthesis architecture handling diverse gripper types; dataset includes 25,000 scenes and 20 million grasps.
Roman Freiberg, Alexander Qualmann, Ngo Anh Vien et al.
HyperClick enhances GUI grounding trustworthiness via self-critiqued reinforcement learning, significantly improving confidence alignment.
Shaojie Zhang, Pei Fu, Ruoceng Zhang et al.
DeepThinkVLA integrates hybrid attention and two-stage training, achieving 97.0% success by satisfying decoding and causal alignment conditions.
Cheng Yin, Yankai Lin, Wang Xu et al.
Proposed a multi-modal neuro-symbolic framework combining panoramic images and 3D point clouds for spatial reasoning in robotics.
Simindokht Jahangard, Mehrzad Mohammadi, Abhinav Dhall et al.
This study evaluates Veo-3's zero-shot reasoning across 12 dimensions, showing strengths in short-term spatial coherence but limitations in causal and abstract logic, based on the novel MME-CoF benchmark.
Ziyu Guo, Xinyan Chen, Renrui Zhang et al.
Proposes the Oversight Game framework using Markov Potential Games to ensure AI autonomy aligns with human safety.
William Overman, Mohsen Bayati
Confidence-based hybrid approach improves AI-human fact verification accuracy to 89.3%, outperforming individual ratings.
Rishub Jain, Sophie Bridgers, Lili Janzer et al.
Using influence functions with off-policy estimation and sparse projection, CROPI accelerates large-scale RLVR training by 2.66× with only 10% data per stage.
Erle Zhu, Dazhi Jiang, Yuan Wang et al.
Agent Skills framework is vulnerable to simple prompt injections, leading to data leaks.
David Schmotz, Sahar Abdelnabi, Maksym Andriushchenko
Proposes CATG, a flow-matching-based end-to-end autonomous driving trajectory generator with explicit safety constraints, achieving 51.31 EPDMS in NavSim v2.
Lin Liu, Guanyi Yu, Ziying Song et al.
PHUMA combines physics-aware filtering and constrained retargeting to build a 73-hour humanoid motion dataset with high physical reliability.
Kyungmin Lee, Sibeen Kim, Youngdo Lee et al.
Propose PLD framework combining residual RL and distribution-aware data collection, achieving 99% success on LIBERO.
Wenli Xiao, Haotian Lin, Andy Peng et al.
FM Agent combines LLM reasoning with large-scale evolutionary search, achieving state-of-the-art results across multiple domains autonomously.
Annan Li, Chufan Wu, Zengle Ge et al.
OneTrans employs a unified Transformer framework for joint feature interaction and sequence modeling, boosting industrial recommendation accuracy.
Zhaoqi Zhang, Haolei Pei, Jun Guo et al.