Data-Efficient RLVR via Off-Policy Influence Guidance
Using influence functions with off-policy estimation and sparse projection, CROPI accelerates large-scale RLVR training by 2.66× with only 10% data per stage.
Erle Zhu, Dazhi Jiang, Yuan Wang et al.