Inferring Capabilities from Task Performance with Bayesian Triangulation
Fuente:
arXiv
Saved in:
| Main Authors: | Burden, John, Voudouris, Konstantinos, Burnell, Ryan, Rutar, Danaja, Cheke, Lucy, Hernández-Orallo, José |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
Predictable Artificial Intelligence
by: Zhou, Lexin, et al.
Published: (2023)
by: Zhou, Lexin, et al.
Published: (2023)
Bringing Comparative Cognition To Computers
by: Voudouris, Konstantinos, et al.
Published: (2025)
by: Voudouris, Konstantinos, et al.
Published: (2025)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
The Animal-AI Environment: A Virtual Laboratory For Comparative Cognition and Artificial Intelligence Research
by: Voudouris, Konstantinos, et al.
Published: (2023)
by: Voudouris, Konstantinos, et al.
Published: (2023)
Conversational Complexity for Assessing Risk in Large Language Models
by: Burden, John, et al.
Published: (2024)
by: Burden, John, et al.
Published: (2024)
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
by: Mecattaf, Matteo G., et al.
Published: (2024)
by: Mecattaf, Matteo G., et al.
Published: (2024)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
by: Voudouris, Konstantinos, et al.
Published: (2026)
by: Voudouris, Konstantinos, et al.
Published: (2026)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
by: Fitz, Stephen, et al.
Published: (2025)
by: Fitz, Stephen, et al.
Published: (2025)
Learning Alternative Ways of Performing a Task
by: Nieves, David, et al.
Published: (2024)
by: Nieves, David, et al.
Published: (2024)
Evaluating AI Evaluation: Perils and Prospects
by: Burden, John
Published: (2024)
by: Burden, John
Published: (2024)
Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025)
by: Testini, Irene, et al.
Published: (2025)
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Visuospatial Perspective Taking in Multimodal Language Models
by: Prunty, Jonathan, et al.
Published: (2026)
by: Prunty, Jonathan, et al.
Published: (2026)
What should an AI assessor optimise for?
by: Romero-Alvarado, Daniel, et al.
Published: (2025)
by: Romero-Alvarado, Daniel, et al.
Published: (2025)
Beyond the high score: Prosocial ability profiles of multi-agent populations
by: Tesic, Marko, et al.
Published: (2025)
by: Tesic, Marko, et al.
Published: (2025)
Framing the Game: How Context Shapes LLM Decision-Making
by: Robinson, Isaac, et al.
Published: (2025)
by: Robinson, Isaac, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Bayesian Networks, Markov Networks, Moralisation, Triangulation: a Categorical Perspective
by: Lorenzin, Antonio, et al.
Published: (2025)
by: Lorenzin, Antonio, et al.
Published: (2025)
Exploiting LLMs' Reasoning Capability to Infer Implicit Concepts in Legal Information Retrieval
by: Nguyen, Hai-Long, et al.
Published: (2024)
by: Nguyen, Hai-Long, et al.
Published: (2024)
Exploring Major Transitions in the Evolution of Biological Cognition With Artificial Neural Networks
by: Voudouris, Konstantinos, et al.
Published: (2025)
by: Voudouris, Konstantinos, et al.
Published: (2025)
Capabilities Ain't All You Need: Measuring Propensities in AI
by: Romero-Alvarado, Daniel, et al.
Published: (2026)
by: Romero-Alvarado, Daniel, et al.
Published: (2026)
Inferring Implicit Goals Across Differing Task Models
by: Tulli, Silvia, et al.
Published: (2025)
by: Tulli, Silvia, et al.
Published: (2025)
Identifying, Evaluating, and Mitigating Risks of AI Thought Partnerships
by: Oktar, Kerem, et al.
Published: (2025)
by: Oktar, Kerem, et al.
Published: (2025)
GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks
by: Spencer, Ryan, et al.
Published: (2025)
by: Spencer, Ryan, et al.
Published: (2025)
From Abstract to Actionable: Pairwise Shapley Values for Explainable AI
by: Xu, Jiaxin, et al.
Published: (2025)
by: Xu, Jiaxin, et al.
Published: (2025)
Discovering Novel LLM Experts via Task-Capability Coevolution
by: Dai, Andrew, et al.
Published: (2026)
by: Dai, Andrew, et al.
Published: (2026)
Can AI Tools Transform Low-Demand Math Tasks? An Evaluation of Task Modification Capabilities
by: Fox, Danielle S., et al.
Published: (2026)
by: Fox, Danielle S., et al.
Published: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Evaluating General-Purpose AI with Psychometrics
by: Wang, Xiting, et al.
Published: (2023)
by: Wang, Xiting, et al.
Published: (2023)
Robust Weighted Triangulation of Causal Effects Under Model Uncertainty
by: Bhattacharya, Rohit, et al.
Published: (2026)
by: Bhattacharya, Rohit, et al.
Published: (2026)
Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
by: Marín, Javier
Published: (2025)
by: Marín, Javier
Published: (2025)
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
by: Ming, Xiaoyang, et al.
Published: (2026)
by: Ming, Xiaoyang, et al.
Published: (2026)
TReF-6: Inferring Task-Relevant Frames from a Single Demonstration for One-Shot Skill Generalization
by: Ding, Yuxuan, et al.
Published: (2025)
by: Ding, Yuxuan, et al.
Published: (2025)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
by: Wong, Annie, et al.
Published: (2025)
by: Wong, Annie, et al.
Published: (2025)
Similar Items
-
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025) -
Predictable Artificial Intelligence
by: Zhou, Lexin, et al.
Published: (2023) -
Bringing Comparative Cognition To Computers
by: Voudouris, Konstantinos, et al.
Published: (2025) -
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)