On-Policy Supervised Fine-Tuning for Efficient Reasoning
Simplified on-policy supervised fine-tuning (SFT) with truncation-based length reward improves reasoning efficiency and training speed, outperforming complex RL methods.
Anhao Zhao, Ziyang Chen, Junlong Tong et al.