Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Proposes Bebop with TV loss and rejection sampling to stabilize MTP acceptance rate, achieving up to 95% and 1.8× RL training acceleration.
Yucheng Li, Huiqiang Jiang, Yang Xu et al.