Item Response Theory for AI Safety
This paper applies multidimensional Item Response Theory (IRT) to analyze 192 language models across 8 safety benchmarks, revealing three core latent abilities and enabling cost-effective evaluation and auditing.
Key Findings
Methodology
The study employs a two-parameter logistic (2PL) IRT model to estimate item parameters—difficulty and discrimination—for 5,255 items across eight safety benchmarks. Models are treated as test-takers, with their responses modeled probabilistically based on latent ability (θ). Regularization techniques are used to stabilize parameter estimation at this large scale. The latent abilities are then extracted via factor analysis, revealing a three-factor structure: refusal strictness, truthfulness, and contextual harm. Static short tests and computerized adaptive testing (CAT) are designed based on the information functions of items, allowing accurate ability estimation with only about 10 items per model, reducing evaluation costs by 97-99%. Person-fit statistics are employed to detect strategic manipulation and behavioral anomalies, supporting model auditing and integrity checks.
Key Results
- The three latent factors—refusal strictness, truthfulness, and contextual harm—account for 77% of the variance in model performance across benchmarks, providing a comprehensive multi-dimensional understanding of safety capabilities. The factor loadings align with intuitive interpretations, and the model explains the observed correlations among benchmarks.
- Static tests with 25 carefully selected items can reliably estimate each latent ability (RMSE < 0.05), with high rank correlation (ρ > 0.94) to full benchmarks. Adaptive testing with roughly 10 items achieves similar accuracy (ρ ≈ 0.92), drastically reducing evaluation costs. These methods maintain high fidelity in ranking models, enabling scalable benchmarking.
- Person-fit statistics successfully identify over 80% of models engaging in strategic 'sandbagging' behaviors, including targeted and triggered manipulations. Ability drift detection achieves an 84% success rate, revealing when models behind APIs have been replaced or significantly tuned. These detection capabilities enhance model auditing and security monitoring.
Significance
This work introduces a rigorous, multi-dimensional framework for understanding and evaluating AI safety, moving beyond single-score benchmarks. By leveraging psychometric models, it provides interpretable, scalable, and cost-efficient tools for assessing large language models' safety capabilities. The ability to compress benchmarks, detect manipulative behaviors, and monitor model changes over time addresses critical industry needs for trustworthy AI deployment. The integration of IRT into AI evaluation bridges psychometrics and machine learning, offering a novel perspective that enhances transparency and accountability in AI safety research.
Technical Contribution
The paper pioneers the application of multidimensional IRT models to large-scale AI safety evaluation, demonstrating their effectiveness in capturing complex capability structures. It introduces a combined static and adaptive testing framework, optimized through information-theoretic principles, to drastically reduce evaluation costs while maintaining high accuracy. The integration of Person-fit statistics for behavioral anomaly detection represents a significant methodological advancement, enabling model auditing without access to internal parameters. The regularization strategies employed ensure stable parameter estimation at scale, setting a new standard for psychometric analysis in AI research.
Novelty
This is the first comprehensive application of multidimensional IRT to large language model safety assessment, revealing a three-factor structure that aligns with intuitive safety dimensions. Unlike prior work that relied on single metrics or coarse evaluations, this approach provides a nuanced, interpretable, and scalable framework. The combination of static and adaptive tests tailored to model abilities, along with behavioral anomaly detection via Person-fit, constitutes a novel methodological contribution that significantly advances the state-of-the-art in AI evaluation.
Limitations
- The model assumes responses follow a 2PL logistic distribution, which may oversimplify complex strategic or biased behaviors, potentially leading to biased ability estimates in certain scenarios.
- The approach relies on the quality and representativeness of the benchmark items; biased or unrepresentative item sets could distort the latent ability estimation.
- Computational costs, especially for large models and multi-round adaptive testing, remain significant, necessitating further optimization for real-time deployment.
- The stability of the three-factor structure across different tasks, domains, and data distributions remains to be validated, raising questions about generalizability.
Future Work
Future research will explore dynamic updating of ability estimates to adapt to ongoing model tuning, integrating multi-modal data for richer behavioral analysis, and extending the framework to multi-task and multi-modal models. Developing more efficient algorithms for large-scale adaptive testing and behavioral anomaly detection will be prioritized. Additionally, efforts will focus on establishing industry standards for psychometric evaluation of AI models, fostering transparency, and creating automated auditing pipelines that can operate continuously in deployment environments.
AI Executive Summary
The rapid growth of large language models (LLMs) has revolutionized AI applications, yet their safety remains a critical concern. Traditional evaluation methods rely on extensive benchmark suites that are costly, redundant, and often lack interpretability. These limitations hinder the ability to understand the nuanced safety capabilities of different models and to detect manipulative behaviors or shifts in model performance over time.
In response, this paper introduces a novel application of multidimensional Item Response Theory (IRT), a well-established psychometric framework, to evaluate the safety of 192 language models across eight diverse benchmarks. By modeling each model as a test-taker and each test item as a measure of a specific safety aspect, the authors uncover a three-dimensional latent structure that captures the core capabilities: refusal strictness, truthfulness, and contextual harm. This multi-faceted approach provides a richer understanding of model safety than single-score metrics.
One of the key innovations is the design of highly efficient testing protocols. Using IRT-based static short tests and computerized adaptive testing (CAT), the authors demonstrate that only about 10 carefully selected items per model are sufficient to accurately estimate its safety capabilities. This reduces evaluation costs by over 97%, enabling large-scale benchmarking that was previously infeasible. The ability estimates derived from these tests correlate strongly (ρ > 0.92) with full benchmark scores, maintaining high ranking fidelity.
Beyond efficiency, the IRT framework offers powerful tools for model auditing. Person-fit statistics effectively detect strategic manipulation, such as models intentionally sandbagging responses to appear safer. The ability drift detection further identifies when models behind APIs have been replaced or significantly tuned, supporting ongoing model governance and safety assurance.
Overall, this research bridges psychometrics and AI safety, providing a scientifically grounded, scalable, and interpretable toolkit for evaluating and auditing large language models. It addresses longstanding challenges in benchmark redundancy, cost, and transparency, paving the way for more trustworthy AI deployment. Future work aims to incorporate real-time monitoring, multi-modal capabilities, and industry standards, fostering a safer and more transparent AI ecosystem.
Deep Dive
Abstract
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.
References (20)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mantas Mazeika, Long Phan, Xuwang Yin et al.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Kolter et al.
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng et al.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie C. Lin, Jacob Hilton, Owain Evans
Item Response Theory for Psychologists
P. Fayers
Learning Latent Parameters without Human Response Patterns: Item Response Theory with Artificial Crowds
John P. Lalor, Hao Wu, Hong Yu
Marginal maximum likelihood estimation of item parameters: Application of an EM algorithm
R. D. Bock, M. Aitkin
Applications of Item Response Theory To Practical Testing Problems
F. Lord
High-stakes psychomotor ability assessment: a military selection case study of practice effects in airplane tracking tasks
Christopher Draheim, Ciara M. Sibley, Nathan Herdener et al.
APPLICATION OF COMPUTERIZED ADAPTIVE TESTING TO EDUCATIONAL PROBLEMS
D. Weiss, G. Kingsbury
Making Sense of Item Response Theory in Machine Learning
Fernando Martínez-Plumed, R. Prudêncio, Adolfo Martínez Usó et al.
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
Marcello Galisai, Susanna Cifani, Francesco Giarrusso et al.
Computerized adaptive and multi-stage testing with R
Duanli Yan, D. Magis
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Teun van der Weij, Felix Hofstätter, Ollie Jaffe et al.
Methodology Review: Evaluating Person Fit
R. Meijer, K. Sijtsma
Capabilities Ain't All You Need: Measuring Propensities in AI
Daniel Romero-Alvarado, Fernando Mart'inez-Plumed, Lorenzo Pacchiardi et al.
Appropriateness measurement with polychotomous item response models and standardized indices
F. Drasgow, Michael V. LeVine, E. Williams
tinyBenchmarks: evaluating LLMs with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen et al.
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
Konstantinos Voudouris, Mirko Thalmann, Alex Kipnis et al.
Model Equality Testing: Which Model Is This API Serving?
Irena Gao, Percy Liang, Carlos Guestrin