Latent On-Policy Self-Distillation
LOPD introduces learnable latent privileged context, outperforming OPSD variants with less than 30% rollout budget.
Guibin Zhang, Jiayang Lyu, Ran Sun et al.
LOPD introduces learnable latent privileged context, outperforming OPSD variants with less than 30% rollout budget.
Guibin Zhang, Jiayang Lyu, Ran Sun et al.
Introduces Action-Conditioned Predictive Consistency (ACPC) to diagnose robustness of JEPA world models against visual perturbations, with theoretical bounds and empirical validation.
Guo An, Zijing Wu, Honghua Dong et al.
Decoupled Contrastive Decoding (DCD) accelerates generation by using an expert-aligned lightweight proposer and only applying amateurs in verification, achieving 1.65-1.95× speedup.
Zhixuan Liu, Zhichen Dong, Yuanfu Wang et al.
ReFind is a keyword-based conversational search method achieving 58.2% accuracy without semantic indexing.
Ruizhe Li, Licheng Zhang, Benfeng Xu et al.
SkillMisevo-Gym reveals skill misevolution in LLMs; SafeEvolve reduces unsafe retrieval by 26.7 percentage points.
Xutao Mao, Liangjie Zhao, Xiang Zheng et al.
RippleNet employs local differential signals with multi-scale, multi-directional modeling and frequency-guided attention to improve cross-generator fake image detection, achieving 89.0% accuracy.
Jiazhen Yang, Ruijin Jin, Junjun Zheng et al.
HybridSB-MoE combines dual-domain Schrödinger bridges with scene-adaptive expert routing, outperforming baselines on VoiceBank+DEMAND with theoretical guarantees.
Zhengyi Lu, Aswini Sivakumar, Jie Hu et al.
Introduces WMRL, replacing environment execution with a world model, accelerating AutoResearch agent training 3-4×, surpassing standard RL.
Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.
StateFlow employs a structured 3D world state model with three stages—construction, evolution, and access—to enable high-quality, controllable previsualization for film and game design.
Yuyang Yin, Zixiang Li, Longxuan Deng et al.
This study introduces an automated KG-DML construction framework using RAG and LLMs for complex system diagnostics.
Saman Marandi, Yu-Shu Hu, Mohammad Modarres
Proposes a formal reward design framework combining objectives, causal features, and preference queries, ensuring conflict-free linear reward parameters via geometric optimization.
Di Yang Shi, W. Bradley Knox
The 'Agentic Self-Improvement' framework employs a two-stage goal-driven optimization, significantly improving semantic adherence and control in image-to-video generation, outperforming unguided methods with up to 69% preference.
Aman Tyagi, Hemanth Boinpally, Jonathan Chen et al.
This paper identifies four structural barriers—web presence gap, token scarcity, tokenization penalty, connectivity exclusion—that hinder AI support for underrepresented languages like Bengali.
Avijit Roy, Proma Roy
Decision-contract theory combines CRC, capacity κ, and actionability; 90.3% of configurations met risk targets with 83.4% correct automation.
Zhenpeng Li
This study reveals simulator collapse in multi-agent RL, proposing Verbalized Sampling and Co-Training to enhance generalization, improving success rates by up to 14%.
Simon Yu, Nicholas Tomlin, Marwa Abdulhai et al.
HAMP-LIC introduces Hessian-based mixed-precision post-training quantization, achieving up to 4.85× compression with only 0.59% BD-rate loss.
Yuefeng Zhang
Proposes an efficient near-optimal adversarial m-set bandit algorithm with high-probability regret bound of O(√dT log(K/δ)), avoiding exponential complexity.
Francesco Bacchiocchi, Tommaso Cesari, Roberto Colomboni
ScreenShot employs a hierarchical transformer pretrained on 40 drug screening datasets, enabling few-shot prediction and active learning to significantly improve combination drug screening efficiency.
Antoine de Mathelin, Christopher Tosh, Wesley Tansey
GAS framework uses decoupled Transformer architecture with NEP for visual understanding, achieving zero inference overhead and significant improvements in spatial perception.
Zhongbin Guo, Jiahao Xie, Dongling Xiao et al.
This paper introduces budget-dependent evaluation of LLMs, revealing model rank reversals across token budgets (64-4096 tokens), with 14.1% of oracle gap captured by a budget-aware router.
Rodrigo Guedes de Souza, Alison R. Panisson