Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Bean, Andrew M., Kearns, Ryan Othniel, Romanou, Angelika, Hafner, Franziska Sofia, Mayne, Harry, Batzner, Jan, Foroutan, Negar, Schmitz, Chris, Korgul, Karolina, Batra, Hunar, Deb, Oishi, Beharry, Emma, Emde, Cornelius, Foster, Thomas, Gausen, Anna, Grandury, María, Han, Simeng, Hofmann, Valentin, Ibrahim, Lujain, Kim, Hazel, Kirk, Hannah Rose, Lin, Fangru, Liu, Gabrielle Kaili-May, Luettgau, Lennart, Magomere, Jabez, Rystrøm, Jonathan, Sotnikova, Anna, Yang, Yushi, Zhao, Yilun, Bibi, Adel, Bosselut, Antoine, Clark, Ronald, Cohan, Arman, Foerster, Jakob, Gal, Yarin, Hale, Scott A., Raji, Inioluwa Deborah, Summerfield, Christopher, Torr, Philip H. S., Ududec, Cozmin, Rocher, Luc, Mahdi, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Agent Benchmarks Fail Public Sector Requirements
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025)
by: Dubois, Magda, et al.
Published: (2025)
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
Oversight Structures for Agentic AI in Public-Sector Organizations
by: Schmitz, Chris, et al.
Published: (2025)
by: Schmitz, Chris, et al.
Published: (2025)
AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
by: Barghorn, Jérémy, et al.
Published: (2026)
by: Barghorn, Jérémy, et al.
Published: (2026)
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
by: Gausen, Anna, et al.
Published: (2026)
by: Gausen, Anna, et al.
Published: (2026)
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
by: Summerfield, Christopher, et al.
Published: (2025)
by: Summerfield, Christopher, et al.
Published: (2025)
Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison
by: de Chillaz, Aymeric, et al.
Published: (2025)
by: de Chillaz, Aymeric, et al.
Published: (2025)
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
A Framework for Exploring the Consequences of AI-Mediated Enterprise Knowledge Access and Identifying Risks to Workers
by: Gausen, Anna, et al.
Published: (2023)
by: Gausen, Anna, et al.
Published: (2023)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
by: Lugoloobi, William, et al.
Published: (2026)
by: Lugoloobi, William, et al.
Published: (2026)
Characterizing and modeling harms from interactions with design patterns in AI interfaces
by: Ibrahim, Lujain, et al.
Published: (2024)
by: Ibrahim, Lujain, et al.
Published: (2024)
When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
by: Khouja, Jude, et al.
Published: (2025)
by: Khouja, Jude, et al.
Published: (2025)
Training language models to be warm and empathetic makes them less reliable and more sycophantic
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024)
by: Bean, Andrew M., et al.
Published: (2024)
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios
by: Mai, Kimberly T., et al.
Published: (2026)
by: Mai, Kimberly T., et al.
Published: (2026)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
by: Mensah, Samuel, et al.
Published: (2025)
by: Mensah, Samuel, et al.
Published: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025)
by: Bhagwatkar, Rishika, et al.
Published: (2025)
The #Somos600M Project: Generating NLP resources that represent the diversity of the languages from LATAM, the Caribbean, and Spain
by: Grandury, María
Published: (2024)
by: Grandury, María
Published: (2024)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
When Do LLM Preferences Predict Downstream Behavior?
by: Slama, Katarina, et al.
Published: (2026)
by: Slama, Katarina, et al.
Published: (2026)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
On a new order of Eocene mammals
by: Marsh, Othniel Charles
Published: (1875)
by: Marsh, Othniel Charles
Published: (1875)
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026)
by: Kearns, Ryan Othniel
Published: (2026)
It's the same but not the same: Do LLMs distinguish Spanish varieties?
by: Mayor-Rocher, Marina, et al.
Published: (2025)
by: Mayor-Rocher, Marina, et al.
Published: (2025)
Vulnerability-Amplifying Interaction Loops: a systematic failure mode in AI chatbot mental-health interactions
by: Weilnhammer, Veith, et al.
Published: (2026)
by: Weilnhammer, Veith, et al.
Published: (2026)
EVCL: Elastic Variational Continual Learning with Weight Consolidation
by: Batra, Hunar, et al.
Published: (2024)
by: Batra, Hunar, et al.
Published: (2024)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
by: Mayor-Rocher, Marina, et al.
Published: (2024)
by: Mayor-Rocher, Marina, et al.
Published: (2024)
One-shot emergency psychiatric triage across 15 frontier AI chatbots
by: Weilnhammer, Veith, et al.
Published: (2026)
by: Weilnhammer, Veith, et al.
Published: (2026)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Similar Items
-
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025) -
Agent Benchmarks Fail Public Sector Requirements
by: Rystrøm, Jonathan, et al.
Published: (2026) -
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025) -
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
by: Magomere, Jabez, et al.
Published: (2025)