cs.AI 2602.13407

On-Policy Supervised Fine-Tuning for Efficient Reasoning

Simplified on-policy supervised fine-tuning (SFT) with truncation-based length reward improves reasoning efficiency and training speed, outperforming complex RL methods.

Anhao Zhao, Ziyang Chen, Junlong Tong et al.

2026-02-14 48
stat.ML 2602.13362

Nonparametric Distribution Regression Re-calibration

Proposes a nonparametric distribution calibration method based on conditional kernel mean embeddings, significantly improving calibration accuracy.

Ádám Jung, Domokos M. Kelen, András A. Benczúr

2026-02-13 49
cs.CL 2602.12275

On-Policy Context Distillation for Language Models

Proposes On-Policy Context Distillation (OPCD), minimizing reverse KL divergence on self-generated trajectories, boosting knowledge internalization for tasks like math reasoning and transfer.

Tianzhu Ye, Li Dong, Xun Wu et al.

2026-02-13 134 citations 44
cs.LG 2602.12233

Categorical Flow Maps

Proposes Categorical Flow Maps, leveraging continuous paths for rapid few-step categorical data generation with state-of-the-art results.

Daan Roos, Oscar Davis, Floor Eijkelboom et al.

2026-02-13 52