cs.LG 2601.13566

Self-Improvement as Coherence Optimization: A Theoretical Account

Proposes coherence optimization as a unified framework for self-improvement, proving its equivalence to description-length regularization, and introduces Gibbs sampling for scalable optimization.

Tianyi Qiu, Ahmed Hani Ismail, Zhonghao He et al.

2026-01-20 37
cs.LG 2601.11516

Building Production-Ready Probes For Gemini

Proposes new long-context robust activation probes, improving misuse detection accuracy and efficiency.

János Kramár, Joshua Engels, Zheng Wang et al.

2026-01-17 40
cs.LG 2512.25070

Scaling Open-Ended Reasoning to Predict the Future

Using OpenForesight dataset, retrieval, and reinforcement learning, trained OpenForecaster 8B achieves competitive future prediction accuracy.

Nikhil Chandak, Shashwat Goel, Ameya Prabhu et al.

2026-01-01 40
cs.LG 2512.23707

Training AI Co-Scientists Using Rubric Rewards

Proposes a rubric-guided reinforcement learning framework to enhance AI research plan generation, achieving 70% expert preference and 12-22% improvements across domains.

Shashwat Goel, Rishi Hazra, Dulhan Jayalath et al.

2025-12-30 29