SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
SCAN: Self-Denoising Monte Carlo Annotation reduces noise in synthetic data, boosting PRM F1 from 19.9 to 59.1 with only 6% inference cost.
Yuyang Ding, Xinyu Shi, Juntao Li et al.