cs.LG 2603.01223

Learn Hard Problems During RL with Reference Guided Fine-tuning

ReGFT leverages human reference solutions to synthesize positive trajectories, significantly alleviating reward sparsity and boosting RL-based mathematical reasoning performance.

Yangzhen Wu, Shanda Li, Zixin Wen et al.

2026-03-02 6 citations 52
cs.LG 2603.13277

Learning Retrieval Models with Sparse Autoencoders

SPLARE introduces a sparse latent retrieval model using pretrained SAEs, outperforming vocabulary-based methods in multilingual and out-of-domain tasks.

Thibault Formal, Maxime Louis, Hervé Dejean et al.

2026-02-27 43
cs.LG 2602.23413

EvoX: Meta-Evolution for Automated Discovery

EvoX employs meta-evolution of search strategies, outperforming AlphaEvolve on nearly 200 tasks with significant efficiency gains.

Shu Liu, Shubham Agarwal, Monishwaran Maheswaran et al.

2026-02-27 42