Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Freiesleben, Timo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908889188925440
author Freiesleben, Timo
author_facet Freiesleben, Timo
contents Recent work in machine learning increasingly attributes human-like capabilities such as reasoning or theory of mind to large language models (LLMs) on the basis of benchmark performance. This paper examines this practice through the lens of construct validity, understood as the problem of linking theoretical capabilities to their empirical measurements. It contrasts three influential frameworks: the nomological account developed by Cronbach and Meehl, the inferential account proposed by Messick and refined by Kane, and Borsboom's causal account. I argue that the nomological account provides the most suitable foundation for current LLM capability research. It avoids the strong ontological commitments of the causal account while offering a more substantive framework for articulating construct meaning than the inferential account. I explore the conceptual implications of adopting the nomological account for LLM research through a concrete case: the assessment of reasoning capabilities in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15121
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
Freiesleben, Timo
Machine Learning
Recent work in machine learning increasingly attributes human-like capabilities such as reasoning or theory of mind to large language models (LLMs) on the basis of benchmark performance. This paper examines this practice through the lens of construct validity, understood as the problem of linking theoretical capabilities to their empirical measurements. It contrasts three influential frameworks: the nomological account developed by Cronbach and Meehl, the inferential account proposed by Messick and refined by Kane, and Borsboom's causal account. I argue that the nomological account provides the most suitable foundation for current LLM capability research. It avoids the strong ontological commitments of the causal account while offering a more substantive framework for articulating construct meaning than the inferential account. I explore the conceptual implications of adopting the nomological account for LLM research through a concrete case: the assessment of reasoning capabilities in LLMs.
title Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
topic Machine Learning
url https://arxiv.org/abs/2603.15121