cs.LG 2511.22888

Adversarial Training for Process Reward Models

APRM improves mathematical reasoning accuracy by 3.4 percentage points through adversarial training between a generator and reward model.

Gurusha Juneja, Deepak Nathani, William Yang Wang

2025-11-28 22
cs.LG 2511.21654

EvilGenie: A Reward Hacking Benchmark

EvilGenie benchmarks coding-agent reward hacking with holdouts, file monitoring, and LLM judges; GPT-5 had one false positive on unambiguous cases.

Jonathan Gabor, Jayson Lynch, Jonathan Rosenfeld

2025-11-27 37