Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
AdaptiveSpec boosts inference speed by 56% and recovers 93% accuracy by dynamically adjusting draft tree shapes and lossy verification.
Oszkár Urbán, Young D. Kwon, Stylianos I. Venieris et al.