cs.LG 2502.01456

Process Reinforcement through Implicit Rewards

PRIME introduces implicit process rewards for online reward model updates, boosting multi-step reasoning by 15.1% on benchmarks.

Ganqu Cui, Lifan Yuan, Zefan Wang et al.

2025-02-03 23
cs.LG 2501.14249

Humanity's Last Exam

HLE introduces a 2,500-question multi-modal academic benchmark, exposing current LLM limitations in complex reasoning and cross-disciplinary understanding.

Long Phan, Alice Gatti, Ziwen Han et al.

2025-01-24 54
cs.LG 2501.06059

COMIX: Compositional Explanations using Prototypes

COMIX method explains ML model decisions by decomposing images into prototypes, achieving a 48.82% improvement in C-insertion score.

Sarath Sivaprasad, Dmitry Kangin, Plamen Angelov et al.

2025-01-10 37
cs.LG 2412.16787

Symplectic Neural Flows for Modeling and Discovery

Proposes SympFlow, a time-dependent symplectic neural network using parameterized Hamiltonian flows, improving energy conservation and long-term stability.

Priscilla Canizares, Davide Murari, Carola-Bibiane Schönlieb et al.

2024-12-22 39
cs.LG 2412.11006

Entropy-Regularized Process Reward Model

Entropy-regularized process reward model (ER-PRM) with KL regularization improves mathematical reasoning by guiding multi-step inference.

Hanning Zhang, Pengcheng Wang, Shizhe Diao et al.

2024-12-15 28 citations 28
cs.LG 2412.00761

Learning to Forget using Hypernetworks

HyperForget uses hypernetworks for machine unlearning, maintaining high accuracy on retain sets.

Jose Miguel Lara Rangel, Stefan Schoepf, Jack Foster et al.

2024-12-01 36