cs.LG 2608.03447

Approximate Speculative Decoding

Proposes Approximate Speculative Decoding (ASD), a zero-training verifier that boosts throughput by up to 15.26% via budgeted longest-prefix selection.

Yuannuo Feng, Zegang Peng, Yuxin Xie et al.

2026-08-04 41
cs.LG 2608.02870

Maglev: Sliding Recurrent Memory

Maglev combines sliding-window attention with fixed-size recurrent memory via a parallel-trained predictor, boosting long-sequence modeling by 1-2% accuracy.

Bo Liu, Qiang Liu

2026-08-04 48
cs.LG 2607.28022

Flux-OPD: On-Policy Distillation with Evolving Contexts

Flux-OPD uses evolving contexts and reverse KL decomposition to stabilize open-domain distillation, outperforming existing methods with +2-3 score improvements.

Yuran Wang, Zekun Wang, Bohan Zeng et al.

2026-07-30 38