General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Lexin, Pacchiardi, Lorenzo, Martínez-Plumed, Fernando, Collins, Katherine M., Moros-Daval, Yael, Zhang, Seraphina, Zhao, Qinlin, Huang, Yitian, Sun, Luning, Prunty, Jonathan E., Li, Zongqian, Sánchez-García, Pablo, Chen, Kexin Jiang, Casares, Pablo A. M., Zu, Jiyun, Burden, John, Mehrbakhsh, Behzad, Stillwell, David, Cebrian, Manuel, Wang, Jindong, Henderson, Peter, Wu, Sherry Tongshuang, Kyllonen, Patrick C., Cheke, Lucy, Xie, Xing, Hernández-Orallo, José |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictable Artificial Intelligence
by: Zhou, Lexin, et al.
Published: (2023)
by: Zhou, Lexin, et al.
Published: (2023)
Capabilities Ain't All You Need: Measuring Propensities in AI
by: Romero-Alvarado, Daniel, et al.
Published: (2026)
by: Romero-Alvarado, Daniel, et al.
Published: (2026)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
Visuospatial Perspective Taking in Multimodal Language Models
by: Prunty, Jonathan, et al.
Published: (2026)
by: Prunty, Jonathan, et al.
Published: (2026)
Conversational Complexity for Assessing Risk in Large Language Models
by: Burden, John, et al.
Published: (2024)
by: Burden, John, et al.
Published: (2024)
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
From Human-Level AI Tales to AI Leveling Human Scales
by: Romero, Peter, et al.
Published: (2026)
by: Romero, Peter, et al.
Published: (2026)
Inferring Capabilities from Task Performance with Bayesian Triangulation
by: Burden, John, et al.
Published: (2023)
by: Burden, John, et al.
Published: (2023)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025)
by: Testini, Irene, et al.
Published: (2025)
Evaluating General-Purpose AI with Psychometrics
by: Wang, Xiting, et al.
Published: (2023)
by: Wang, Xiting, et al.
Published: (2023)
What should an AI assessor optimise for?
by: Romero-Alvarado, Daniel, et al.
Published: (2025)
by: Romero-Alvarado, Daniel, et al.
Published: (2025)
Multi-agent AI systems outperform human teams in creativity
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
by: Mecattaf, Matteo G., et al.
Published: (2024)
by: Mecattaf, Matteo G., et al.
Published: (2024)
Additional 18th century records of endemic Seychelles fauna.
by: Cheke, Anthony
Published: (2008)
by: Cheke, Anthony
Published: (2008)
Amor a Distancia: Nuevas Formas de Vida en la Era Global
by: Marta Plumed
Published: (2013)
by: Marta Plumed
Published: (2013)
Peter J. Jones 1945–2024
by: Robert A. Cheke
Published: (2025)
by: Robert A. Cheke
Published: (2025)
PromptBench: A Unified Library for Evaluation of Large Language Models
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
Dynamic Evaluation of Large Language Models by Meta Probing Agents
by: Zhu, Kaijie, et al.
Published: (2024)
by: Zhu, Kaijie, et al.
Published: (2024)
Planificación turística, promoción y sostenibilidad ambiental: el caso de España
by: Marta Plumed Lasarte
Published: (2018)
by: Marta Plumed Lasarte
Published: (2018)
Incremental and developmental perspectives for general-purpose learning systems
by: Fernando Martínez-Plumed
Published: (2017)
by: Fernando Martínez-Plumed
Published: (2017)
Análisis de las variables influyentes en la aceptación de una estrategia de branding territorial por parte de la población local. El caso de Zaragoza (España)
by: Marta Plumed Lasarte
Published: (2017)
by: Marta Plumed Lasarte
Published: (2017)
Conversions explicites entre des fonctions sommatoires de la fonction de Möbius
by: Daval, Florian
Published: (2020)
by: Daval, Florian
Published: (2020)
Sur la somme de Möbius $\sum_{n \leqslant x} μ(n)n^{-s}$ autour de $s=1$ et des sommes dérivées, première étude
by: Daval, Florian
Published: (2024)
by: Daval, Florian
Published: (2024)
Design, Synthesis, Computational Modeling and Biological Evaluation of Novel N‐( 7 H ‐Pyrrolo[2,3‐ d ]pyrimidin‐4‐yl)cinnamamides as Potential Covalent Inhibitors for Oncology
by: Rameshwar S. Cheke, et al.
Published: (2025)
by: Rameshwar S. Cheke, et al.
Published: (2025)
Mathematics and its history / John Stillwell
by: Stillwell, John
by: Stillwell, John
Characterization of Brassolaeliocattleya Raye Holmes ‘Mendenhall’ - putatively transformed for resistance to Cymbidium mosaic virus
by: Nyan Stillwell
Published: (2013)
by: Nyan Stillwell
Published: (2013)
Augmenting Rating-Scale Measures with Text-Derived Items Using the Information-Determined Scoring (IDS) Framework
by: Watson, Joe, et al.
Published: (2025)
by: Watson, Joe, et al.
Published: (2025)
Andrew D. Berns, The Bible and natural philosophy in Renaissance Italy: Jewish and Christian physicians in search of truth, Nueva York, Cambridge University Press, 2015, 309 páginas
by: Jesús de Prado Plumed
Published: (2016)
by: Jesús de Prado Plumed
Published: (2016)
Sarissa Carneiro, Retórica del infortunio. Persuasión, deleite y ejemplaridad en el siglo XVI, Madrid/Frankfurt, Iberoamericana/Vervuert, 2015, 235 páginas
by: Jesús de Prado Plumed
Published: (2017)
by: Jesús de Prado Plumed
Published: (2017)
The treatment of madness in the nineteenth and twentieth centuries: discourses about curability in Spanish mental health care, 1890-1917
by: José Javier Plumed Domingo
Published: (2016)
by: José Javier Plumed Domingo
Published: (2016)
The Future of Scholarly Communication. Minutes of the Meeting of the Association of Research Libraries (95th, Washington, D.C., October 17-18, 1979).
by: Daval, Nicola, Ed.
Published: (1980)
by: Daval, Nicola, Ed.
Published: (1980)
Resources for Research Libraries. Minutes of the Meeting of the Association of Research Libraries (98th, New York, New York, May 7-8, 1981).
by: Daval, Nicola, Ed.
Published: (1981)
by: Daval, Nicola, Ed.
Published: (1981)
ARL: Setting the Agenda for the 1990s. Minutes of the [Membership] Meeting (112th, Oakland, California, May 5-6, 1988).
by: Daval, Nicola, Ed.
Published: (1989)
by: Daval, Nicola, Ed.
Published: (1989)
Government Information in Electronic Format. Minutes of the Meeting of the Association of Research Libraries (110th, Pittsburgh, Pennsylvania, May 7-8, 1987).
by: Daval, Nicola, Ed.
Published: (1987)
by: Daval, Nicola, Ed.
Published: (1987)
The Restrictive Effects of Government Information Policies on Scholarship and Research. Minutes of the Meeting of the Association of Research Libraries (107th, October 23-24, 1985, Washington, D.C.).
by: Daval, Nicola, Ed.
Published: (1986)
by: Daval, Nicola, Ed.
Published: (1986)
Research Libraries: Measurement, Management, Marketing. Minutes of the Meeting of the Association of Research Libraries (108th, Minneapolis, Minnesota, May 1-2, 1986).
by: Daval, Nicola, Ed.
Published: (1986)
by: Daval, Nicola, Ed.
Published: (1986)
Similar Items
-
Predictable Artificial Intelligence
by: Zhou, Lexin, et al.
Published: (2023) -
Capabilities Ain't All You Need: Measuring Propensities in AI
by: Romero-Alvarado, Daniel, et al.
Published: (2026) -
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)