The Evaluation Trap: Benchmark Design as Theoretical Commitment
Fuente:
arXiv
Saved in:
| Main Author: | Kalaitzidis, Theodore J |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI
by: O'Herlihy, Michael, et al.
Published: (2026)
by: O'Herlihy, Michael, et al.
Published: (2026)
Do AI Companies Make Good on Voluntary Commitments to the White House?
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Designing AI-Resilient Assessments Using Interconnected Problems: A Theoretically Grounded and Empirically Validated Framework
by: Ding, Kaihua
Published: (2025)
by: Ding, Kaihua
Published: (2025)
The Controllability Trap: A Governance Framework for Military AI Agents
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
Administrative Law's Fourth Settlement: AI and the Capability-Accountability Trap
by: Caputo, Nicholas
Published: (2026)
by: Caputo, Nicholas
Published: (2026)
The Reasoning Under Uncertainty Trap: A Structural AI Risk
by: Pilditch, Toby D.
Published: (2024)
by: Pilditch, Toby D.
Published: (2024)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026)
by: Li, Zongjie, et al.
Published: (2026)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
by: Shi, Shaojie, et al.
Published: (2026)
by: Shi, Shaojie, et al.
Published: (2026)
LegalScore: Development of a Benchmark for Evaluating AI Models in Legal Career Exams in Brazil
by: Caparroz, Roberto, et al.
Published: (2025)
by: Caparroz, Roberto, et al.
Published: (2025)
A Multimodal Manufacturing Safety Chatbot: Knowledge Base Design, Benchmark Development, and Evaluation of Multiple RAG Approaches
by: Singh, Ryan, et al.
Published: (2025)
by: Singh, Ryan, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Designing an Interdisciplinary Artificial Intelligence Curriculum for Engineering: Evaluation and Insights from Experts
by: Schleiss, Johannes, et al.
Published: (2025)
by: Schleiss, Johannes, et al.
Published: (2025)
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
by: Munir, Sheza, et al.
Published: (2026)
by: Munir, Sheza, et al.
Published: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
The Alignment Trap: Complexity Barriers
by: Yao, Jasper
Published: (2025)
by: Yao, Jasper
Published: (2025)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems
by: Dai, Gordon, et al.
Published: (2025)
by: Dai, Gordon, et al.
Published: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Machine Learning Techniques with Fairness for Prediction of Completion of Drug and Alcohol Rehabilitation
by: Roberts-Licklider, Karen, et al.
Published: (2024)
by: Roberts-Licklider, Karen, et al.
Published: (2024)
LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation
by: Lee, Yong Oh, et al.
Published: (2025)
by: Lee, Yong Oh, et al.
Published: (2025)
LLM Psychosis: A Theoretical and Diagnostic Framework for Reality-Boundary Failures in Large Language Models
by: Raj, Ashutosh
Published: (2026)
by: Raj, Ashutosh
Published: (2026)
Co-Designing Interdisciplinary Design Projects with AI
by: Liow, Wei Ting, et al.
Published: (2025)
by: Liow, Wei Ting, et al.
Published: (2025)
Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
by: Bakumenko, Alexander, et al.
Published: (2025)
by: Bakumenko, Alexander, et al.
Published: (2025)
Agent Benchmarks Fail Public Sector Requirements
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
by: Li, Miles Q., et al.
Published: (2026)
by: Li, Miles Q., et al.
Published: (2026)
Evaluating AI Evaluation: Perils and Prospects
by: Burden, John
Published: (2024)
by: Burden, John
Published: (2024)
The Non-Optimality of Scientific Knowledge: Path Dependence, Lock-In, and The Local Minimum Trap
by: Mabrok, Mohamed
Published: (2026)
by: Mabrok, Mohamed
Published: (2026)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
by: Shen, Judy Hanwen, et al.
Published: (2025)
by: Shen, Judy Hanwen, et al.
Published: (2025)
Benchmarking Large Language Models on Homework Assessment in Circuit Analysis
by: Chen, Liangliang, et al.
Published: (2025)
by: Chen, Liangliang, et al.
Published: (2025)
S$^3$IT: A Benchmark for Spatially Situated Social Intelligence Test
by: Sun, Zhe, et al.
Published: (2025)
by: Sun, Zhe, et al.
Published: (2025)
Critically Engaged Pragmatism: A Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools
by: Lee, Carole J.
Published: (2026)
by: Lee, Carole J.
Published: (2026)
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
TRIAGE: Ethical Benchmarking of AI Models Through Mass Casualty Simulations
by: Kirch, Nathalie Maria, et al.
Published: (2024)
by: Kirch, Nathalie Maria, et al.
Published: (2024)
TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS
by: Anlen, Shirin, et al.
Published: (2025)
by: Anlen, Shirin, et al.
Published: (2025)
The Emotional Alignment Design Policy
by: Schwitzgebel, Eric, et al.
Published: (2025)
by: Schwitzgebel, Eric, et al.
Published: (2025)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Similar Items
-
Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI
by: O'Herlihy, Michael, et al.
Published: (2026) -
Do AI Companies Make Good on Voluntary Commitments to the White House?
by: Wang, Jennifer, et al.
Published: (2025) -
Designing AI-Resilient Assessments Using Interconnected Problems: A Theoretically Grounded and Empirically Validated Framework
by: Ding, Kaihua
Published: (2025) -
The Controllability Trap: A Governance Framework for Military AI Agents
by: Sahoo, Subramanyam
Published: (2026) -
Administrative Law's Fourth Settlement: AI and the Capability-Accountability Trap
by: Caputo, Nicholas
Published: (2026)