cs.AI 2608.05086

Item Response Theory for AI Safety

This paper applies multidimensional Item Response Theory (IRT) to analyze 192 language models across 8 safety benchmarks, revealing three core latent abilities and enabling cost-effective evaluation and auditing.

Joshua Fonseca Rivera, Neil Shah, David Demitri Africa et al.

2026-08-06 127
cs.AI 2608.01735

DAPD: Dual-Anchored Policy Distillation

DAPD significantly improves privilege illusion by dual-path and dual-source anchoring, with an average gain of 2.00 points.

Jianyu Wu, Yizhou Wang, Encheng Su et al.

2026-08-03 2