Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
Fuente:
arXiv
Saved in:
| Main Authors: | Pacchiardi, Lorenzo, Tesic, Marko, Cheke, Lucy G., Hernández-Orallo, José |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025)
by: Testini, Irene, et al.
Published: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
by: Mecattaf, Matteo G., et al.
Published: (2024)
by: Mecattaf, Matteo G., et al.
Published: (2024)
HANS, are you clever? Clever Hans Effect Analysis of Neural Systems
by: Ranaldi, Leonardo, et al.
Published: (2023)
by: Ranaldi, Leonardo, et al.
Published: (2023)
Beyond the high score: Prosocial ability profiles of multi-agent populations
by: Tesic, Marko, et al.
Published: (2025)
by: Tesic, Marko, et al.
Published: (2025)
Inferring Capabilities from Task Performance with Bayesian Triangulation
by: Burden, John, et al.
Published: (2023)
by: Burden, John, et al.
Published: (2023)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
The Clever Hans Effect in Unsupervised Learning
by: Kauffmann, Jacob, et al.
Published: (2024)
by: Kauffmann, Jacob, et al.
Published: (2024)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
Visuospatial Perspective Taking in Multimodal Language Models
by: Prunty, Jonathan, et al.
Published: (2026)
by: Prunty, Jonathan, et al.
Published: (2026)
Conversational Complexity for Assessing Risk in Large Language Models
by: Burden, John, et al.
Published: (2024)
by: Burden, John, et al.
Published: (2024)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
by: Fitz, Stephen, et al.
Published: (2025)
by: Fitz, Stephen, et al.
Published: (2025)
Bringing Comparative Cognition To Computers
by: Voudouris, Konstantinos, et al.
Published: (2025)
by: Voudouris, Konstantinos, et al.
Published: (2025)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
keqing: knowledge-based question answering is a nature chain-of-thought mentor of LLM
by: Wang, Chaojie, et al.
Published: (2023)
by: Wang, Chaojie, et al.
Published: (2023)
From Clever Hans to Scientific Discovery: Interpreting EEG Foundational Transformers with LRP
by: Bexten, Justus Meyer zu, et al.
Published: (2026)
by: Bexten, Justus Meyer zu, et al.
Published: (2026)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
by: V, Venktesh, et al.
Published: (2024)
by: V, Venktesh, et al.
Published: (2024)
Dynamic benchmarking framework for LLM-based conversational data capture
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
LLMzSzŁ: a comprehensive LLM benchmark for Polish
by: Jassem, Krzysztof, et al.
Published: (2025)
by: Jassem, Krzysztof, et al.
Published: (2025)
A benchmark for joint dialogue satisfaction, emotion recognition, and emotion state transition prediction
by: Bian, Jing, et al.
Published: (2026)
by: Bian, Jing, et al.
Published: (2026)
KeyKnowledgeRAG (K^2RAG): An Enhanced RAG method for improved LLM question-answering capabilities
by: Markondapatnaikuni, Hruday, et al.
Published: (2025)
by: Markondapatnaikuni, Hruday, et al.
Published: (2025)
Multi-agent AI systems outperform human teams in creativity
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Seven simple steps for log analysis in AI systems
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
Hypothetical answers to continuous queries over data streams
by: Cruz-Filipe, Luís, et al.
Published: (2019)
by: Cruz-Filipe, Luís, et al.
Published: (2019)
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation
by: Govindarajan, Hariprasath, et al.
Published: (2025)
by: Govindarajan, Hariprasath, et al.
Published: (2025)
Measuring Spurious Correlation in Classification: 'Clever Hans' in Translationese
by: Borah, Angana, et al.
Published: (2023)
by: Borah, Angana, et al.
Published: (2023)
Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction
by: Saito, Kuniaki, et al.
Published: (2024)
by: Saito, Kuniaki, et al.
Published: (2024)
On the effectiveness of LLMs for automatic grading of open-ended questions in Spanish
by: Capdehourat, Germán, et al.
Published: (2025)
by: Capdehourat, Germán, et al.
Published: (2025)
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
by: Maharjan, Jenish, et al.
Published: (2024)
by: Maharjan, Jenish, et al.
Published: (2024)
FormationEval, an open multiple-choice benchmark for petroleum geoscience
by: Ermilov, Almaz
Published: (2026)
by: Ermilov, Almaz
Published: (2026)
Suvach -- Generated Hindi QA benchmark
by: Narayanan, Vaishak, et al.
Published: (2024)
by: Narayanan, Vaishak, et al.
Published: (2024)
Question answering systems for health professionals at the point of care -- a systematic review
by: Kell, Gregory, et al.
Published: (2024)
by: Kell, Gregory, et al.
Published: (2024)
Agribot: agriculture-specific question answer system
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
From text to multimodal: a survey of adversarial example generation in question answering systems
by: Yigit, Gulsum, et al.
Published: (2023)
by: Yigit, Gulsum, et al.
Published: (2023)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
Your Large Language Models Are Leaving Fingerprints
by: McGovern, Hope, et al.
Published: (2024)
by: McGovern, Hope, et al.
Published: (2024)
ACL-Verbatim: hallucination-free question answering for research
by: Recski, Gábor, et al.
Published: (2026)
by: Recski, Gábor, et al.
Published: (2026)
Similar Items
-
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024) -
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
by: Burden, John, et al.
Published: (2025) -
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025) -
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025) -
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
by: Mecattaf, Matteo G., et al.
Published: (2024)