cs.LG 2511.22888

Adversarial Training for Process Reward Models

APRM improves mathematical reasoning accuracy by 3.4 percentage points through adversarial training between a generator and reward model.

Gurusha Juneja, Deepak Nathani, William Yang Wang

2025-11-28 3
cs.LG 2511.21654

EvilGenie: A Reward Hacking Benchmark

EvilGenie benchmarks coding-agent reward hacking with holdouts, file monitoring, and LLM judges; GPT-5 had one false positive on unambiguous cases.

Jonathan Gabor, Jayson Lynch, Jonathan Rosenfeld

2025-11-27 23
cs.LG 2511.17339

ReBaPL: Repulsive Bayesian Prompt Learning

ReBaPL integrates cyclical SGHMC and representation-space repulsion to enhance multi-modal prompt Bayesian inference, capturing multi-peak posteriors for better generalization.

Yassir Bendou, Omar Ezzahir, Eduardo Fernandes Montesuma et al.

2025-11-22 58
cs.LG 2511.08094

Stuart-Landau Oscillatory Graph Neural Network

Stuart-Landau Oscillatory Graph Neural Network (SLGNN) dynamically adjusts amplitudes to address oversmoothing in GNNs.

Kaicheng Zhang, David N. Reynolds, Piero Deidda et al.

2025-11-11 9