Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
Fuente:
arXiv
Saved in:
| Main Author: | Freiesleben, Timo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
by: Freiesleben, Timo, et al.
Published: (2025)
by: Freiesleben, Timo, et al.
Published: (2025)
Artificial Neural Nets and the Representation of Human Concepts
by: Freiesleben, Timo
Published: (2023)
by: Freiesleben, Timo
Published: (2023)
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026)
by: Freiesleben, Timo, et al.
Published: (2026)
Performative Validity of Recourse Explanations
by: König, Gunnar, et al.
Published: (2025)
by: König, Gunnar, et al.
Published: (2025)
Improvement-Focused Causal Recourse (ICR)
by: König, Gunnar, et al.
Published: (2022)
by: König, Gunnar, et al.
Published: (2022)
Scientific Inference With Interpretable Machine Learning: Analyzing Models to Learn About Real-World Phenomena
by: Freiesleben, Timo, et al.
Published: (2022)
by: Freiesleben, Timo, et al.
Published: (2022)
CountARFactuals -- Generating plausible model-agnostic counterfactual explanations with adversarial random forests
by: Dandl, Susanne, et al.
Published: (2024)
by: Dandl, Susanne, et al.
Published: (2024)
[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games
by: Pollo, Jorge Carrasco, et al.
Published: (2026)
by: Pollo, Jorge Carrasco, et al.
Published: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
Phase-Adaptive LLM Framework with Multi-Stage Validation for Construction Robot Task Allocation: A Systematic Benchmark Against Traditional Optimization Algorithms
by: Kaitha, Shyam prasad reddy, et al.
Published: (2025)
by: Kaitha, Shyam prasad reddy, et al.
Published: (2025)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks
by: Srivastava, Ashutosh, et al.
Published: (2025)
by: Srivastava, Ashutosh, et al.
Published: (2025)
Benchmarking of Clustering Validity Measures Revisited
by: Simpson, Connor, et al.
Published: (2025)
by: Simpson, Connor, et al.
Published: (2025)
Reassessing the Validity of Spurious Correlations Benchmarks
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
IGL-Bench: Establishing the Comprehensive Benchmark for Imbalanced Graph Learning
by: Qin, Jiawen, et al.
Published: (2024)
by: Qin, Jiawen, et al.
Published: (2024)
Discrete Semantic States and Hamiltonian Dynamics in LLM Embedding Spaces
by: Laine, Timo Aukusti
Published: (2025)
by: Laine, Timo Aukusti
Published: (2025)
PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data
by: Cheng, Pu, et al.
Published: (2026)
by: Cheng, Pu, et al.
Published: (2026)
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
by: Zhou, Yu, et al.
Published: (2024)
by: Zhou, Yu, et al.
Published: (2024)
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents
by: Pleines, Marco, et al.
Published: (2023)
by: Pleines, Marco, et al.
Published: (2023)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
QUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability
by: Zou, Bo, et al.
Published: (2026)
by: Zou, Bo, et al.
Published: (2026)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
by: Macina, Jakub, et al.
Published: (2025)
by: Macina, Jakub, et al.
Published: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation
by: Han, Andy, et al.
Published: (2026)
by: Han, Andy, et al.
Published: (2026)
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
by: Anand, Ashwani, et al.
Published: (2026)
by: Anand, Ashwani, et al.
Published: (2026)
Approximation Capabilities of Feedforward Neural Networks with GELU Activations
by: Yakovlev, Konstantin, et al.
Published: (2025)
by: Yakovlev, Konstantin, et al.
Published: (2025)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
by: Skelic, Lejla, et al.
Published: (2025)
by: Skelic, Lejla, et al.
Published: (2025)
A Modular Dataset to Demonstrate LLM Abstraction Capability
by: Atanas, Adam, et al.
Published: (2025)
by: Atanas, Adam, et al.
Published: (2025)
Generating Statistical Charts with Validation-Driven LLM Workflows
by: Poličar, Pavlin G., et al.
Published: (2026)
by: Poličar, Pavlin G., et al.
Published: (2026)
Selective Adversarial Attacks on LLM Benchmarks
by: Dubrovsky, Ivan, et al.
Published: (2025)
by: Dubrovsky, Ivan, et al.
Published: (2025)
Exploiting Interpretable Capabilities with Concept-Enhanced Diffusion and Prototype Networks
by: Carballo-Castro, Alba, et al.
Published: (2024)
by: Carballo-Castro, Alba, et al.
Published: (2024)
Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
by: Macar, Uzay, et al.
Published: (2025)
by: Macar, Uzay, et al.
Published: (2025)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
by: Bhowmik, Shimanto, et al.
Published: (2025)
by: Bhowmik, Shimanto, et al.
Published: (2025)
AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
by: Shi, Yichen, et al.
Published: (2025)
by: Shi, Yichen, et al.
Published: (2025)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025)
by: Dong, Peijie, et al.
Published: (2025)
Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
by: Wan, Weilin, et al.
Published: (2025)
by: Wan, Weilin, et al.
Published: (2025)
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
Similar Items
-
The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
by: Freiesleben, Timo, et al.
Published: (2025) -
Artificial Neural Nets and the Representation of Human Concepts
by: Freiesleben, Timo
Published: (2023) -
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026) -
Performative Validity of Recourse Explanations
by: König, Gunnar, et al.
Published: (2025) -
Improvement-Focused Causal Recourse (ICR)
by: König, Gunnar, et al.
Published: (2022)