Reinforced Planning with Latent World Models
RP1 learns multi-step plan improvement rules via reinforcement, achieving near-perfect success with 1/1000 model rollouts and 67× faster inference in robotic tasks.
Armin Sommer, Jannik Schilling
RP1 learns multi-step plan improvement rules via reinforcement, achieving near-perfect success with 1/1000 model rollouts and 67× faster inference in robotic tasks.
Armin Sommer, Jannik Schilling
LabDex evaluates dexterous lab manipulation through hierarchical tasks, validating its effectiveness in real and simulated environments.
Zhipeng Tang, Sihang Chen, Sha Zhang et al.
DAEI combines residual denoising autoencoder with generative inversion, achieving 154% BLEU improvement in noisy text embedding inversion.
Yubo Wang, Shujie Cui, James Bailey et al.
VA-Judger employs chain-of-thought reasoning and multi-dimensional reinforcement learning to align video-audio generation with human preferences, outperforming metric-based rewards.
Yinming Huang, Shuyuan Tu, Xi Yan et al.
OneModel unifies multi-scenario ranking using long-context user modeling, achieving +0.33% Time Spent and +8.18% CTR online.
Yinqi Zhang, Peiyu Hu, Yuntian Tang et al.
Introducing a analytical grid cell model combined with boundary vector cells reduces spatial aliasing by 94-99%, validated across three environments.
Alexander Johnson, Obadah Ghizawi, Ali A. Minai
WhiteMatter dynamically mixes all layer states into KV channels; 16-layer full-cache PPL reaches 19.968, 8.2% below vanilla.
Wenbo Zhang, Xiang Ren
OmniAlign unifies multilingual word and sentence alignment, excelling in long-text scenarios.
Mengpeng Yang, Jingxu Yang, Chao Chen et al.
Developed a multimodal RAG system using ColQwen2-VL-2b-instruct and GPT-4.1, achieving 93.37% recall@5 for Cessna 172 manual retrieval.
Seongjun Ha, Md Rashedul Islam, Gaurav Nanda et al.
Fine-tuning poetry data improves Arabic idiom understanding by 2.33%; cultural fine-tuning shows no significant effect.
Mena Attia, Mona Diab, Thamar Solorio
Task-conditioned least-privilege framework reduces excess permissions to 0.79%, boosting safe success to 98.48% on 1500 tasks.
Alexander Tu, Michael Tu
GAPL integrates LLM estimation, simulation grounding, and PPO optimization to improve trajectory planning, reducing collision rate by 0.76 and increasing reward by 1.44.
Zhihong Cui, Hengyu Liu, Zhangkai Wu et al.
Proposes a capability-centric data infrastructure with multi-stage curriculum scheduling, curating 440M images for generalist image generation, enabling effective multi-capability transfer.
Xingjian Wang, Zhao Wang, Taihang Hu et al.
This study systematically evaluates the fragility of memory-based self-improving agents, revealing high variance and task order sensitivity, and proposes information enrichment to mitigate performance degradation.
Qinyuan Ye, Yu Li, Yada Pruksachatkun et al.
Using latent variable graded response models, this study measures users' willingness to deploy and accept agent-mediated communication in online dating, revealing a significant asymmetry with deployment willingness three times higher.
Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev et al.
HLSR integrates real-time edge speeds with short-term forecasts, using dual-threshold congestion detection and driver personalization, significantly improving congestion mitigation efficiency.
Xiao Wang, Shun Ren Yang, Hui Nien Hung
Proposes Bayesian Optimization-based sampling schedule (OYS) that reduces steps to 5 while retaining 89%-94% of quality, greatly lowering inference costs.
Travis Zhang, Christian Belardi, Justin Lovelace et al.
Proposes a plug-and-play traffic element awareness framework that significantly improves end-to-end autonomous driving across various models, with detailed 3D detection and topological encoding.
Zongzheng Zhang, Jijun Wang, Saining Zhang et al.
Proposes the Effectiveness–Losslessness Framework for tokenization, emphasizing coordinate systems in symbolic music for improved predictive compression.
Yi Wang
GS-Voxel: Fitting-free structured latents for large-scale 3DGS generation, supporting millions of voxels.
Ming Qian, Zijian Wang, Minchao Sun et al.