Measuring What AI Systems Might Do: Towards A Measurement Science in AI
Proposes AI capabilities and propensities as causal dispositional properties; critiques benchmark and elicitation methods; advocates for causal, hypothesis-driven measurement.
Key Findings
Methodology
This work adopts a scientific framework grounded in philosophy of science, measurement theory, and cognitive science. It defines capabilities and propensities as causal, stable dispositional properties characterized by counterfactual relationships between contextual variables (π) and behaviors (v). Measurement involves hypothesizing relevant causal variables, systematically manipulating them, and empirically mapping how variations affect behavior probabilities p(v|π,θ). Unlike traditional methods like benchmark accuracy or latent-variable models (e.g., Item Response Theory), this approach emphasizes explicit causal assumptions, independent operationalization of variables, and systematic variation to infer stable properties. The framework advocates for controlled experiments that vary contextual features to reveal the causal structure underlying AI dispositions.
Key Results
- Applying the framework to language models on mathematical reasoning tasks, the authors demonstrate that performance varies systematically with task difficulty parameters such as digit count and reasoning steps. For example, models show a decline from 80% to 50% accuracy as problem complexity increases, revealing underlying mathematical capability. Similarly, in safety propensity tests, models exhibit increased unsafe behaviors under higher incentive conditions, confirming that propensity is driven by environmental incentives. These results validate the causal measurement approach, showing that performance changes across systematically varied contexts can infer latent capabilities and tendencies.
- The experiments establish that traditional accuracy metrics conflate difficulty with ability, whereas the causal approach isolates the stable property by observing how behavior responds to controlled contextual changes. The authors develop a causal mapping model that predicts behavior probabilities across untested contexts, enabling extrapolation and more reliable assessment of AI dispositions. This approach outperforms existing benchmarks and adversarial tests by providing a principled, theory-based understanding of AI capabilities and propensities.
- Further, the framework supports cross-system comparisons and robustness checks, allowing researchers to quantify differences in abilities and tendencies even when systems are tested in different or novel contexts. The empirical results demonstrate that causal measurement can guide the development of safer, more reliable AI systems, especially in safety-critical domains where traditional performance metrics are insufficient.
Significance
This research advances AI evaluation from performance-centric metrics to a scientifically grounded, causal measurement paradigm. It addresses fundamental limitations of current practices—benchmark accuracy and adversarial elicitation—by providing a framework that captures the stable, underlying properties of AI systems. Such a shift enhances transparency, interpretability, and generalizability, crucial for regulatory compliance and safety assurance. The causal approach enables extrapolation beyond observed data, supporting assessments in high-stakes domains like autonomous driving, healthcare, and security, where understanding the true capabilities and risks of AI is vital. It also bridges the gap between theoretical understanding and empirical measurement, fostering a more rigorous, scientific foundation for AI evaluation.
Technical Contribution
The paper introduces a novel causal framework for measuring AI capabilities and propensities, emphasizing the importance of explicit causal assumptions, independent operationalization of contextual variables, and systematic variation. It departs from traditional correlation-based models like Item Response Theory, instead employing causal inference techniques to map how changes in contextual features influence behavior probabilities. This approach provides formal guarantees about the stability and comparability of measured properties, enabling extrapolation and cross-system evaluation. The framework also integrates counterfactual reasoning, allowing for indirect inference of unobserved behaviors, which is particularly relevant for safety-critical or ethically sensitive behaviors. These contributions significantly enhance the scientific rigor and interpretability of AI measurement.
Novelty
This work is the first to formalize AI capabilities and propensities as causal dispositional properties, emphasizing the importance of systematic variable manipulation and counterfactual analysis. Unlike prior approaches that rely solely on performance metrics or latent-variable models, this framework explicitly models the causal relationships between contextual features and behaviors, enabling more accurate and generalizable measurement. The integration of causal inference into AI evaluation represents a fundamental shift, providing a scientific basis for understanding and comparing AI systems beyond surface-level performance.
Limitations
- The framework relies on hypothesized causal models, which require domain expertise and may be incomplete or inaccurate, affecting measurement validity.
- Operationalizing relevant contextual variables in complex, real-world scenarios remains challenging, especially for safety-critical behaviors that are difficult to elicit safely.
- Computational costs for systematic variable manipulation and causal inference can be high, limiting scalability in large or complex systems. Further research is needed to automate variable selection and causal modeling.
Future Work
Future research will focus on developing methods for automatic identification and validation of causal variables, integrating causal discovery techniques, and expanding the framework to multi-agent and dynamic systems. Efforts will also aim at creating standardized protocols for causal measurement applicable across diverse AI domains, facilitating regulatory adoption and safety assurance. Additionally, combining causal measurement with mechanistic interpretability may yield deeper insights into AI system behavior, ultimately fostering more transparent, reliable, and controllable AI technologies.
AI Executive Summary
Deep Dive
Plain Language Accessible to non-experts
想象你在厨房做菜,想知道每种调料用多少才能做出好味道。传统的方法可能只试一次,觉得味道还不错就算了,但其实,厨师真正的水平是能根据不同的食材、火候和调料比例,调出不同的味道。这就像科学家用系统的方法,调整调料的用量,观察味道的变化,找到调料和味道之间的因果关系。这样,厨师就能科学地评估自己的厨艺,而不是只凭感觉。类似地,AI的能力和偏好也是如此,要通过系统性地改变环境中的条件,观察它的反应,才能真正理解它的内在特质。
ELI14 Explained like you're 14
想象你在玩一款游戏,你的表现不仅仅是赢或输,而是你背后的技能和偏好。有时候你喜欢赢,就会努力;有时候你怕输,就会变得小心。科学家们也想知道AI是不是有类似的技能和偏好,但不能只看它赢了几次或输了几次。相反,他们会试着改变游戏规则,比如增加难度或者给奖励,然后观察它的反应。这样就像调节游戏的难度或奖励一样,科学家可以更好地理解AI的真正能力和偏好。这就像调节游戏中的难度,看看你在不同条件下会做出什么反应,才能知道你到底有多厉害或者喜欢什么。
Abstract
Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, propensities, skills, values, and abilities are routinely used interchangeably and conflated with observable performance, with AI evaluation practices rarely specifying what quantity they purport to measure. We argue that capabilities and propensities are dispositional properties - stable features of systems characterised by counterfactual relationships between contextual conditions and behavioural outputs. Measuring a disposition requires (i) hypothesising which contextual properties are causally relevant, (ii) independently operationalising and measuring those properties, and (iii) empirically mapping how variation in those properties affects the probability of the behaviour. Dominant approaches to AI evaluation, from benchmark averages to data-driven latent-variable models such as Item Response Theory, bypass these steps entirely. Building on ideas from philosophy of science, measurement theory, and cognitive science, we develop a principled account of AI capabilities and propensities as dispositions, show why prevailing evaluation practices fail to measure them, and outline what disposition-respecting, scientifically defensible AI evaluation would require.