The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Devin, Andrew James
Format: Recurso digital
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901166500085760
author Devin, Andrew James
author_facet Devin, Andrew James
contents <p>Artificial intelligence benchmarks are presented as evidence of capability, safety and progress. This paper demonstrates that they are structurally incapable of detecting the failures that matter. Current evaluation methodologies measure pattern-matching performance within training distributions while remaining categorically blind to the absence of verification capability. Verification is used here in the strict epistemic sense: the architectural authority of a system to categorically halt or refuse output under conditions of irreducible ignorance, rather than probabilistic calibration, internal consistency, counterfactual simulation or evidence-conditioned continuation. This is not a failure of measurement optimisation, but a prior condition: benchmarks cannot measure verification where no verification capability exists, even when evaluation is delegated to more advanced AI models.</p> <p>Analysis of major evaluation frameworks (including MMLU [18], GSM8K [19], HumanEval [20], HellaSwag [21] and TruthfulQA [22]) reveals a systematic conflation of statistical correlation with understanding. More critically, optimisation against these benchmarks actively amplifies the failure modes documented in Artificially Confabulating and Untrustworthy (ACU) systems [24]: benchmark gaming rewards confident completion over epistemic restraint, fluent articulation over verified accuracy and sophisticated pattern exploitation over genuine reasoning.</p> <p>The result is an evaluation paradigm that makes unsafe systems appear safe. High benchmark scores correlate with increased confabulation confidence, not reduced confabulation frequency. This finding has immediate regulatory implications: frameworks that reference benchmark performance as safety evidence institutionalise precisely the wrong incentive structure. No evaluation methodology operating within the current paradigm can measure trustworthiness. Trustworthiness requires verification capability (V), which transformer-based architectures structurally lack [24]. The benchmark crisis is therefore conceptual, not technical. Its resolution requires governance frameworks that assume V = 0 and mandate external verification, not better tests administered to systems incapable of the capability being tested.</p> <p>This paper establishes that no benchmark supplying ground truth can, even in principle, measure verification capability, rendering evaluation-based assurance logically invalid rather than merely empirically unreliable.</p> <p>Summary: The first paper establishes the architectural limit. This paper explains why evaluation cannot detect it.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18223209
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
Devin, Andrew James
AI evaluation, AI Benchmarks, Structural confabulation, Verification capability, Trustworthiness in AI, Transformer architectures, AI governance, Category error in AI evaluation, Epistemic verification, KEV framework, Large language models
<p>Artificial intelligence benchmarks are presented as evidence of capability, safety and progress. This paper demonstrates that they are structurally incapable of detecting the failures that matter. Current evaluation methodologies measure pattern-matching performance within training distributions while remaining categorically blind to the absence of verification capability. Verification is used here in the strict epistemic sense: the architectural authority of a system to categorically halt or refuse output under conditions of irreducible ignorance, rather than probabilistic calibration, internal consistency, counterfactual simulation or evidence-conditioned continuation. This is not a failure of measurement optimisation, but a prior condition: benchmarks cannot measure verification where no verification capability exists, even when evaluation is delegated to more advanced AI models.</p> <p>Analysis of major evaluation frameworks (including MMLU [18], GSM8K [19], HumanEval [20], HellaSwag [21] and TruthfulQA [22]) reveals a systematic conflation of statistical correlation with understanding. More critically, optimisation against these benchmarks actively amplifies the failure modes documented in Artificially Confabulating and Untrustworthy (ACU) systems [24]: benchmark gaming rewards confident completion over epistemic restraint, fluent articulation over verified accuracy and sophisticated pattern exploitation over genuine reasoning.</p> <p>The result is an evaluation paradigm that makes unsafe systems appear safe. High benchmark scores correlate with increased confabulation confidence, not reduced confabulation frequency. This finding has immediate regulatory implications: frameworks that reference benchmark performance as safety evidence institutionalise precisely the wrong incentive structure. No evaluation methodology operating within the current paradigm can measure trustworthiness. Trustworthiness requires verification capability (V), which transformer-based architectures structurally lack [24]. The benchmark crisis is therefore conceptual, not technical. Its resolution requires governance frameworks that assume V = 0 and mandate external verification, not better tests administered to systems incapable of the capability being tested.</p> <p>This paper establishes that no benchmark supplying ground truth can, even in principle, measure verification capability, rendering evaluation-based assurance logically invalid rather than merely empirically unreliable.</p> <p>Summary: The first paper establishes the architectural limit. This paper explains why evaluation cannot detect it.</p>
title The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
topic AI evaluation, AI Benchmarks, Structural confabulation, Verification capability, Trustworthiness in AI, Transformer architectures, AI governance, Category error in AI evaluation, Epistemic verification, KEV framework, Large language models
url https://doi.org/10.5281/zenodo.18223209