stat.ML 2502.03261

CARROT: A Cost Aware Rate Optimal Router

CARROT predicts model cost and accuracy, achieving minimax optimal routing; outperforms baselines on SPROUT dataset.

Seamus Somerstep, Felipe Maia Polo, Allysson Flavio Melo de Oliveira et al.

2025-02-05 32 citations 56
cs.LG 2502.02538

Flow Q-Learning

Flow Q-Learning (FQL) uses flow matching to model complex action distributions, avoiding recursive backpropagation, and achieves state-of-the-art results on 73 tasks.

Seohong Park, Qiyang Li, Sergey Levine

2025-02-05 35
cs.LG 2502.01456

Process Reinforcement through Implicit Rewards

PRIME introduces implicit process rewards for online reward model updates, boosting multi-step reasoning by 15.1% on benchmarks.

Ganqu Cui, Lifan Yuan, Zefan Wang et al.

2025-02-03 28