cs.LG 2602.08489

Beyond Correctness: Learning Robust Reasoning via Transfer

RLTR introduces cross-model transfer rewards to enhance reasoning robustness and efficiency, achieving +3.6% in Maj@64 on MATH-500 and 2.5× training step reduction.

Hyunseok Lee, Soheil Abbasloo, Jihoon Tack et al.

2026-02-09 2 citations 39