Measuring What AI Systems Might Do: Towards A Measurement Science in AI

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Voudouris, Konstantinos, Thalmann, Mirko, Kipnis, Alex, Hernández-Orallo, José, Schulz, Eric
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911474410061824
author Voudouris, Konstantinos
Thalmann, Mirko
Kipnis, Alex
Hernández-Orallo, José
Schulz, Eric
author_facet Voudouris, Konstantinos
Thalmann, Mirko
Kipnis, Alex
Hernández-Orallo, José
Schulz, Eric
contents Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, propensities, skills, values, and abilities are routinely used interchangeably and conflated with observable performance, with AI evaluation practices rarely specifying what quantity they purport to measure. We argue that capabilities and propensities are dispositional properties - stable features of systems characterised by counterfactual relationships between contextual conditions and behavioural outputs. Measuring a disposition requires (i) hypothesising which contextual properties are causally relevant, (ii) independently operationalising and measuring those properties, and (iii) empirically mapping how variation in those properties affects the probability of the behaviour. Dominant approaches to AI evaluation, from benchmark averages to data-driven latent-variable models such as Item Response Theory, bypass these steps entirely. Building on ideas from philosophy of science, measurement theory, and cognitive science, we develop a principled account of AI capabilities and propensities as dispositions, show why prevailing evaluation practices fail to measure them, and outline what disposition-respecting, scientifically defensible AI evaluation would require.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00063
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Measuring What AI Systems Might Do: Towards A Measurement Science in AI
Voudouris, Konstantinos
Thalmann, Mirko
Kipnis, Alex
Hernández-Orallo, José
Schulz, Eric
Computers and Society
Artificial Intelligence
Machine Learning
Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, propensities, skills, values, and abilities are routinely used interchangeably and conflated with observable performance, with AI evaluation practices rarely specifying what quantity they purport to measure. We argue that capabilities and propensities are dispositional properties - stable features of systems characterised by counterfactual relationships between contextual conditions and behavioural outputs. Measuring a disposition requires (i) hypothesising which contextual properties are causally relevant, (ii) independently operationalising and measuring those properties, and (iii) empirically mapping how variation in those properties affects the probability of the behaviour. Dominant approaches to AI evaluation, from benchmark averages to data-driven latent-variable models such as Item Response Theory, bypass these steps entirely. Building on ideas from philosophy of science, measurement theory, and cognitive science, we develop a principled account of AI capabilities and propensities as dispositions, show why prevailing evaluation practices fail to measure them, and outline what disposition-respecting, scientifically defensible AI evaluation would require.
title Measuring What AI Systems Might Do: Towards A Measurement Science in AI
topic Computers and Society
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.00063