HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning
Diffusion Policy enables 79.3% success in multi-object prosthetic grasping via imitation learning.
Carlo Alessi, Federico Vasile, Federico Ceola et al.
Diffusion Policy enables 79.3% success in multi-object prosthetic grasping via imitation learning.
Carlo Alessi, Federico Vasile, Federico Ceola et al.
Proposes multi-layer attention as an amplifier of demonstration effectiveness; uses gradient flow to select demonstrations, achieving 6.8% average improvement.
Dingzirui Wang, Xuangliang Zhang, Keyan Xu et al.
Proposes a Transformer-based unified framework achieving 85% mIoU in multimodal referring segmentation.
Henghui Ding, Song Tang, Shuting He et al.
RecoMind framework optimizes in-session user satisfaction in recommendation systems using RL, boosting video watch time by 15.81%.
Mehdi Ben Ayed, Fei Feng, Jay Adams et al.
XRoboToolkit is a cross-platform XR-based robot teleoperation framework with low-latency stereoscopic feedback and optimized inverse kinematics.
Zhigen Zhao, Liuchuan Yu, Ke Jing et al.
Proposes LLM-based autonomous agents with task planning and tool integration, significantly advancing software automation.
Yihong Dong, Xue Jiang, Jiaru Qian et al.
Seed-Prover achieves 78.1% proof success on IMO problems using formal verification and long chain reasoning.
Luoxin Chen, Jinming Gu, Liankai Huang et al.
Introduced a causal explanation method for concept drift, enhancing model actionability.
David Komnick, Kathrin Lammers, Barbara Hammer et al.
Trae Agent employs modular agents for repository-level issue fixing, achieving Pass@1 of 75.20%, outperforming SOTA by 10.22%.
Trae Research Team, Pengfei Gao, Zhao Tian et al.
Proposes TP-GRPO with process rewards, achieving +4.32% to +6.67% in Pass@1 on math reasoning models with fewer samples.
Tao He, Rongchuan Mu, Lizi Liao et al.
PixNerd introduces an end-to-end pixel diffusion model using neural fields, achieving 2.15 FID on ImageNet 256×256 without VAE or cascades.
Shuai Wang, Ziteng Gao, Chenhui Zhu et al.
Constructed Bengali multilingual benchmarks, evaluated 10 open-source LLMs, found size and tokenization impact performance, with larger models showing more robustness.
Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Sazzad Islam et al.
Introduced Alpha-CLIP and SMS score to improve OV-3DIS, achieving 32.7% mAP on ScanNet200.
Sanghun Jung, Jingjing Zheng, Ke Zhang et al.
ControlMed enhances medical language model efficiency by controlling reasoning length, outperforming existing models in experiments.
Sung-Min Lee, Siyoon Lee, Juyeon Kim et al.
Spec-VLA employs relaxed acceptance speculative decoding to accelerate VLA models, achieving 1.42× speedup with 44% longer acceptance length.
Songsheng Wang, Rucheng Yu, Zhihang Yuan et al.
W-CFM uses Gibbs kernel weighting to approximate Entropic OT, improving path straightness and sample quality efficiently.
Sergio Calvo-Ordonez, Matthieu Meunier, Alvaro Cartea et al.
GRID framework enhances generative recommendation using Semantic IDs, significantly improving Recall@10.
Clark Mingxuan Ju, Liam Collins, Leonardo Neves et al.
Using random matrix theory, the paper compares asymptotic performance of Stack-SVD and SVD-Stack in high-dimensional data integration.
Tavor Z. Baharav, Phillip B. Nicol, Rafael A. Irizarry et al.
X-Omni integrates reinforcement learning into discrete autoregressive models, significantly improving image quality and instruction-following capabilities, enabling unified multimodal generation.
Zigang Geng, Yibing Wang, Yeyao Ma et al.
UserBench evaluates user interaction capabilities, finding models align with user intent only 20% of the time.
Cheng Qian, Zuxin Liu, Akshara Prabhakar et al.