cs.LG 2608.13040

Latent On-Policy Self-Distillation

LOPD introduces learnable latent privileged context, outperforming OPSD variants with less than 30% rollout budget.

Guibin Zhang, Jiayang Lyu, Ran Sun et al.

2026-08-13 70
cs.CL 2608.12913

Decoupled Contrastive Decoding via Expert-Aligned Drafting

Decoupled Contrastive Decoding (DCD) accelerates generation by using an expert-aligned lightweight proposer and only applying amateurs in verification, achieving 1.65-1.95× speedup.

Zhixuan Liu, Zhichen Dong, Yuanfu Wang et al.

2026-08-13 46
cs.LG 2608.12564

Scaling Automatic Research Agents via World Models

Introduces WMRL, replacing environment execution with a world model, accelerating AutoResearch agent training 3-4×, surpassing standard RL.

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.

2026-08-13 58
cs.CV 2608.12290

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

The 'Agentic Self-Improvement' framework employs a two-stage goal-driven optimization, significantly improving semantic adherence and control in image-to-video generation, outperforming unguided methods with up to 69% preference.

Aman Tyagi, Hemanth Boinpally, Jonathan Chen et al.

2026-08-13 110