From Human-Level AI Tales to AI Leveling Human Scales

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Romero, Peter, Martínez-Plumed, Fernando, Tidler, Zachary R., Téhénan, Matthieu, Chen, Sipeng, Antón, Álvaro David Gómez, Sun, Luning, Cebrian, Manuel, Zhou, Lexin, Daval, Yael Moros, Romero-Alvarado, Daniel, Pérez, Félix Martí, Wei, Kevin, Hernández-Orallo, José
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908940485263360
author Romero, Peter
Martínez-Plumed, Fernando
Tidler, Zachary R.
Téhénan, Matthieu
Chen, Sipeng
Antón, Álvaro David Gómez
Sun, Luning
Cebrian, Manuel
Zhou, Lexin
Daval, Yael Moros
Romero-Alvarado, Daniel
Pérez, Félix Martí
Wei, Kevin
Hernández-Orallo, José
author_facet Romero, Peter
Martínez-Plumed, Fernando
Tidler, Zachary R.
Téhénan, Matthieu
Chen, Sipeng
Antón, Álvaro David Gómez
Sun, Luning
Cebrian, Manuel
Zhou, Lexin
Daval, Yael Moros
Romero-Alvarado, Daniel
Pérez, Félix Martí
Wei, Kevin
Hernández-Orallo, José
contents Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world population' and report performance on a common, human-anchored scale. Concretely, we build on a set of multi-level scales for different capabilities where each level should represent a probability of success of the whole world population on a logarithmic scale with a base $B$. We calibrate each scale for each capability (reasoning, comprehension, knowledge, volume, etc.) by compiling publicly released human test data spanning education and reasoning benchmarks (PISA, TIMSS, ICAR, UKBioBank, and ReliabilityBench). The base $B$ is estimated by extrapolating between samples with two demographic profiles using LLMs, with the hypothesis that they condense rich information about human populations. We evaluate the quality of different mappings using group slicing and post-stratification. The new techniques allow for the recalibration and standardization of scales relative to the whole-world population.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18911
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Human-Level AI Tales to AI Leveling Human Scales
Romero, Peter
Martínez-Plumed, Fernando
Tidler, Zachary R.
Téhénan, Matthieu
Chen, Sipeng
Antón, Álvaro David Gómez
Sun, Luning
Cebrian, Manuel
Zhou, Lexin
Daval, Yael Moros
Romero-Alvarado, Daniel
Pérez, Félix Martí
Wei, Kevin
Hernández-Orallo, José
Machine Learning
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world population' and report performance on a common, human-anchored scale. Concretely, we build on a set of multi-level scales for different capabilities where each level should represent a probability of success of the whole world population on a logarithmic scale with a base $B$. We calibrate each scale for each capability (reasoning, comprehension, knowledge, volume, etc.) by compiling publicly released human test data spanning education and reasoning benchmarks (PISA, TIMSS, ICAR, UKBioBank, and ReliabilityBench). The base $B$ is estimated by extrapolating between samples with two demographic profiles using LLMs, with the hypothesis that they condense rich information about human populations. We evaluate the quality of different mappings using group slicing and post-stratification. The new techniques allow for the recalibration and standardization of scales relative to the whole-world population.
title From Human-Level AI Tales to AI Leveling Human Scales
topic Machine Learning
url https://arxiv.org/abs/2602.18911