AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hardy, Michael, Reuel, Anka, Zhang, Lijin, Casabianca, Jodi M., Truong, Sang, Dave, Yash, Lee, Hansol, Domingue, Benjamin, Koyejo, Sanmi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914598842531840
author Hardy, Michael
Reuel, Anka
Zhang, Lijin
Casabianca, Jodi M.
Truong, Sang
Dave, Yash
Lee, Hansol
Domingue, Benjamin
Koyejo, Sanmi
author_facet Hardy, Michael
Reuel, Anka
Zhang, Lijin
Casabianca, Jodi M.
Truong, Sang
Dave, Yash
Lee, Hansol
Domingue, Benjamin
Koyejo, Sanmi
contents While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis (CFA) and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance and establish: (1) structures assumed in current reporting practice underestimate the strength of relationships between benchmarks; (2) evidence of local dependence among leaderboard items, undermining uses of benchmarks as measurement instruments under current scoring systems; (3) contributor metadata explains more rank-relevant variance ($\approx9\%$) than architecture or deployment categories in this context; (4) a manifest-score "scaling law" slope has low reliability ($R_β=0.53$); by contrast, the latent general-factor size slope is highly stable across ecosystem controls ($R_g=0.97$). We are able to provide unique insights into benchmark dynamics, such as which benchmarks are a function of LLM size and which can be oppositely impacted by post-training practices. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25272
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Hardy, Michael
Reuel, Anka
Zhang, Lijin
Casabianca, Jodi M.
Truong, Sang
Dave, Yash
Lee, Hansol
Domingue, Benjamin
Koyejo, Sanmi
Artificial Intelligence
Computers and Society
Applications
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis (CFA) and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance and establish: (1) structures assumed in current reporting practice underestimate the strength of relationships between benchmarks; (2) evidence of local dependence among leaderboard items, undermining uses of benchmarks as measurement instruments under current scoring systems; (3) contributor metadata explains more rank-relevant variance ($\approx9\%$) than architecture or deployment categories in this context; (4) a manifest-score "scaling law" slope has low reliability ($R_β=0.53$); by contrast, the latent general-factor size slope is highly stable across ecosystem controls ($R_g=0.97$). We are able to provide unique insights into benchmark dynamics, such as which benchmarks are a function of LLM size and which can be oppositely impacted by post-training practices. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.
title AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
topic Artificial Intelligence
Computers and Society
Applications
url https://arxiv.org/abs/2605.25272