Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Burden, John, Tešić, Marko, Pacchiardi, Lorenzo, Hernández-Orallo, José |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
par: Testini, Irene, et autres
Publié: (2025)
par: Testini, Irene, et autres
Publié: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
par: Pacchiardi, Lorenzo, et autres
Publié: (2025)
par: Pacchiardi, Lorenzo, et autres
Publié: (2025)
What should an AI assessor optimise for?
par: Romero-Alvarado, Daniel, et autres
Publié: (2025)
par: Romero-Alvarado, Daniel, et autres
Publié: (2025)
Capabilities Ain't All You Need: Measuring Propensities in AI
par: Romero-Alvarado, Daniel, et autres
Publié: (2026)
par: Romero-Alvarado, Daniel, et autres
Publié: (2026)
Beyond the high score: Prosocial ability profiles of multi-agent populations
par: Tesic, Marko, et autres
Publié: (2025)
par: Tesic, Marko, et autres
Publié: (2025)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
par: Rutar, Danaja, et autres
Publié: (2025)
par: Rutar, Danaja, et autres
Publié: (2025)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
par: Voudouris, Konstantinos, et autres
Publié: (2026)
par: Voudouris, Konstantinos, et autres
Publié: (2026)
From Abstract to Actionable: Pairwise Shapley Values for Explainable AI
par: Xu, Jiaxin, et autres
Publié: (2025)
par: Xu, Jiaxin, et autres
Publié: (2025)
Physics-Guided Transformer (PGT): Physics-Aware Attention Mechanism for PINNs
par: Zeraatkar, Ehsan, et autres
Publié: (2026)
par: Zeraatkar, Ehsan, et autres
Publié: (2026)
Conversational Complexity for Assessing Risk in Large Language Models
par: Burden, John, et autres
Publié: (2024)
par: Burden, John, et autres
Publié: (2024)
Seven simple steps for log analysis in AI systems
par: Dubois, Magda, et autres
Publié: (2026)
par: Dubois, Magda, et autres
Publié: (2026)
Learning Paradigms and Modelling Methodologies for Digital Twins in Process Industry
par: Mayr, Michael, et autres
Publié: (2024)
par: Mayr, Michael, et autres
Publié: (2024)
Evaluating AI Evaluation: Perils and Prospects
par: Burden, John
Publié: (2024)
par: Burden, John
Publié: (2024)
MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer
par: Palattuparambil, Ajsal Shereef, et autres
Publié: (2025)
par: Palattuparambil, Ajsal Shereef, et autres
Publié: (2025)
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
par: Panigrahy, Deepak, et autres
Publié: (2026)
par: Panigrahy, Deepak, et autres
Publié: (2026)
Technical Report: Evaluating Goal Drift in Language Model Agents
par: Arike, Rauno, et autres
Publié: (2025)
par: Arike, Rauno, et autres
Publié: (2025)
Evaluating Temporal and Structural Anomaly Detection Paradigms for DDoS Traffic
par: Lima, Yasmin Souza, et autres
Publié: (2026)
par: Lima, Yasmin Souza, et autres
Publié: (2026)
Inferring Capabilities from Task Performance with Bayesian Triangulation
par: Burden, John, et autres
Publié: (2023)
par: Burden, John, et autres
Publié: (2023)
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
par: Liu, Chenruo, et autres
Publié: (2025)
par: Liu, Chenruo, et autres
Publié: (2025)
Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
par: Singh, Sagalpreet, et autres
Publié: (2025)
par: Singh, Sagalpreet, et autres
Publié: (2025)
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
par: Karbevski, Marko
Publié: (2026)
par: Karbevski, Marko
Publié: (2026)
Advancing Software Engineering in the AI-ML Paradigm: A Study of Optimized Methodologies and Obstacles
par: Dr. Rohan Jain and Dr. Leela Rao
Publié: (2023)
par: Dr. Rohan Jain and Dr. Leela Rao
Publié: (2023)
Evaluating the Goal-Directedness of Large Language Models
par: Everitt, Tom, et autres
Publié: (2025)
par: Everitt, Tom, et autres
Publié: (2025)
Dual Goal Representations
par: Park, Seohong, et autres
Publié: (2025)
par: Park, Seohong, et autres
Publié: (2025)
Measuring Goal-Directedness
par: MacDermott, Matt, et autres
Publié: (2024)
par: MacDermott, Matt, et autres
Publié: (2024)
Proposing Hierarchical Goal-Conditioned Policy Planning in Multi-Goal Reinforcement Learning
par: Rens, Gavin B.
Publié: (2025)
par: Rens, Gavin B.
Publié: (2025)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
par: Choi, Jinwoo, et autres
Publié: (2026)
par: Choi, Jinwoo, et autres
Publié: (2026)
Goal Exploration via Adaptive Skill Distribution for Goal-Conditioned Reinforcement Learning
par: Wu, Lisheng, et autres
Publié: (2024)
par: Wu, Lisheng, et autres
Publié: (2024)
Toward Carbon-Neutral Human AI: Rethinking Data, Computation, and Learning Paradigms for Sustainable Intelligence
par: Santosh, KC, et autres
Publié: (2025)
par: Santosh, KC, et autres
Publié: (2025)
AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift
par: Baek, Eunsu, et autres
Publié: (2025)
par: Baek, Eunsu, et autres
Publié: (2025)
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
par: Ferguson, Nick, et autres
Publié: (2026)
par: Ferguson, Nick, et autres
Publié: (2026)
A Decision-driven Methodology for Designing Uncertainty-aware AI Self-Assessment
par: Canal, Gregory, et autres
Publié: (2024)
par: Canal, Gregory, et autres
Publié: (2024)
Adaptive Semantic Token Selection for AI-native Goal-oriented Communications
par: Devoto, Alessio, et autres
Publié: (2024)
par: Devoto, Alessio, et autres
Publié: (2024)
Action-Sufficient Goal Representations
par: Hyeon, Jinu, et autres
Publié: (2026)
par: Hyeon, Jinu, et autres
Publié: (2026)
Goal Recognition as Reinforcement Learning
par: Amado, Leonardo Rosa, et autres
Publié: (2022)
par: Amado, Leonardo Rosa, et autres
Publié: (2022)
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
par: Reuel, Anka, et autres
Publié: (2025)
par: Reuel, Anka, et autres
Publié: (2025)
Towards Measuring Goal-Directedness in AI Systems
par: Xu, Dylan, et autres
Publié: (2024)
par: Xu, Dylan, et autres
Publié: (2024)
Incoherence in Goal-Conditioned Autoregressive Models
par: Karwowski, Jacek, et autres
Publié: (2025)
par: Karwowski, Jacek, et autres
Publié: (2025)
Documents similaires
-
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
par: Pacchiardi, Lorenzo, et autres
Publié: (2024) -
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
par: Pacchiardi, Lorenzo, et autres
Publié: (2024) -
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
par: Testini, Irene, et autres
Publié: (2025) -
PredictaBoard: Benchmarking LLM Score Predictability
par: Pacchiardi, Lorenzo, et autres
Publié: (2025) -
What should an AI assessor optimise for?
par: Romero-Alvarado, Daniel, et autres
Publié: (2025)