cs.LG 2506.17238

Training a Scientific Reasoning Model for Chemistry

Ether0, a 24B-parameter scientific reasoning model trained via reinforcement learning, outperforms existing chemistry models on molecular design tasks with high data efficiency.

Siddharth M. Narayanan, James D. Braza, Ryan-Rhys Griffiths et al.

2025-06-05 41
cs.LG 2506.04178

OpenThoughts: Data Recipes for Reasoning Models

Systematic data recipes improve reasoning models, achieving 53% on AIME 2025, 51% on LiveCodeBench, and 54% on GPQA, surpassing previous models.

Etash Guha, Ryan Marten, Sedrick Keh et al.

2025-06-05 33
cs.LG 2506.00700

Central Path Proximal Policy Optimization

C3PO introduces a central path-inspired modification to PPO, improving constraint satisfaction and reward performance in constrained RL.

Nikola Milosevic, Johannes Müller, Nico Scherf

2025-06-01 27
cs.LG 2505.24492

Object Centric Concept Bottlenecks

Object-Centric Concept Bottlenecks (OCB) integrates pretrained object detection and concept discovery to enhance performance and interpretability in complex visual tasks, achieving 68.84% accuracy on COCOLogic.

David Steinmann, Wolfgang Stammer, Antonia Wüst et al.

2025-05-30 12 citations 50