cs.CY 2601.23112

How Should AI Safety Benchmarks Benchmark Safety?

Proposes risk management-based improvements for AI safety benchmarks, analyzing 210 tools with probabilistic and validity enhancements.

Cheng Yu, Severin Engelmann, Ruoxuan Cao et al.

2026-01-30 43
cs.CL 2601.21968

OVD: On-policy Verbal Distillation

OVD: trajectory matching with verbal scores reduces memory, improves Web QA and math reasoning by up to 25.7%.

Jing Xiong, Hui Shen, Shansan Gong et al.

2026-01-30 56