Robots Need More than VLA and World Models
Proposes four key interfaces to enhance robot intelligence: data, embodiment, world models, rewards.
Elis Karcini, Faisal Mehrban, Quang Nguyen et al.
Proposes four key interfaces to enhance robot intelligence: data, embodiment, world models, rewards.
Elis Karcini, Faisal Mehrban, Quang Nguyen et al.
Proposes Compress-Distill, compresses reasoning traces to 8.6-21%, reduces training tokens by 12-30%, speeds up training 2-7.6×, retains 96% accuracy.
Maxime Griot, Paul Steven Scotti, Tanishq Mathew Abraham
T-FunS3D employs task-driven hierarchical open-vocabulary 3D segmentation, boosting speed and resource efficiency.
Jingkun Feng, Reza Sabzevari
CausalPhys benchmark with expert-annotated causal graphs evaluates VLM causal reasoning, improving accuracy by 30% after CRFT fine-tuning.
Tianyi Tang, Zhuoyi Lin, Zeyu Feng et al.
Proposes biomedical world models learning multiscale latent states and intervention-conditioned dynamics for future simulation and decision support.
Guangyu Wang, Jingkun Yue, Siqi Zhang et al.
Proposes a geometry-aware dataset condensation method to enhance fidelity and distribution coverage in diffusion model training.
Xiao Cui, Yulei Qin, Mo Zhu et al.
SubtleMemory benchmark evaluates AI agents' fine-grained relational memory discrimination, revealing current systems' weaknesses.
Wenxuan Wang, Haoyu Sun, Fukuan Hou et al.
MARDoc employs structured memory to improve multimodal long document QA, achieving 57.1% accuracy and outperforming baselines.
Kaifeng Chen, Hongtao Liu, Qiyao Peng et al.
AdaPLD achieves efficient decoding with adaptive retrieval and reuse, boosting speed by 3.10×.
Runheng Liu, Jincheng Xie, Wen Hu et al.
Parallel Jacobi Decoding achieves 4.8x to 6.4x speedup in autoregressive image generation.
Boya Liao, Ying Li, Siyong Jian et al.
MolE-RAG integrates literature, molecular features, and structural similarity to enhance LLM-based molecular property prediction, boosting ROC-AUC by up to 28% and reducing RMSE by 67%.
Joey Chan, Wonbin Kweon, Ashley Shin et al.
Discrete-WAM employs shared discrete vision-action tokens for unified world-policy modeling, significantly improving autonomous driving planning performance.
Ziyang Yao, Haochen Liu, Yuncheng Jiang et al.
WorldBench constructs a multi-domain visual concept taxonomy, evaluating multimodal models; top accuracy is only 64%, revealing significant gaps.
Yida Yin, Harish Krishnakumar, Chung Peng Lee et al.
SelfBootTok decomposes images into global and local tokens via self-supervised learning, achieving 1.56 gFID with only 64 tokens, surpassing previous methods.
Haozhe Chi, Jinghan Li, Hao Jiang et al.
BloomBench evaluates VLM cognition across six Bloom levels using 7,747 bilingual items and reports 98.45% audited quality.
Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor et al.
SciVisAgentSkills enhances coding agents for scientific data analysis and visualization efficiency.
Kuangshi Ai, Haichao Miao, Kaiyuan Tang et al.
Zeroth-order optimization reveals a single dominant decoding layer for efficient LLM fine-tuning, achieving up to 4.52× speedup.
Wanhao Yu, Ziyan Wang, Zheng Wang et al.
EpiEvolve adapts a frozen epidemic LLM through memory, reaching 0.629 accuracy and cutting post-shift recovery from 5 to 2 weeks.
Yiming Lu, Sihang Zeng, Zhengxu Tang et al.
Introduced α-STOP for early failure alerting in dialogues and LLM-agent trajectories, improving frontier quality by 1-42%.
Avinash Baidya, Xinran Liang, Ruocheng Guo et al.
ALE benchmark evaluates AI on long-term, high-value real-world industry tasks; current top models achieve less than 1% success rate on hardest tasks.
Yiyou Sun, Xinyang Han, Weichen Zhang et al.