The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901166500085760 |
|---|---|
| author | Devin, Andrew James |
| author_facet | Devin, Andrew James |
| contents | <p>Artificial intelligence benchmarks are presented as evidence of capability, safety and progress. This paper demonstrates that they are structurally incapable of detecting the failures that matter. Current evaluation methodologies measure pattern-matching performance within training distributions while remaining categorically blind to the absence of verification capability. Verification is used here in the strict epistemic sense: the architectural authority of a system to categorically halt or refuse output under conditions of irreducible ignorance, rather than probabilistic calibration, internal consistency, counterfactual simulation or evidence-conditioned continuation. This is not a failure of measurement optimisation, but a prior condition: benchmarks cannot measure verification where no verification capability exists, even when evaluation is delegated to more advanced AI models.</p> <p>Analysis of major evaluation frameworks (including MMLU [18], GSM8K [19], HumanEval [20], HellaSwag [21] and TruthfulQA [22]) reveals a systematic conflation of statistical correlation with understanding. More critically, optimisation against these benchmarks actively amplifies the failure modes documented in Artificially Confabulating and Untrustworthy (ACU) systems [24]: benchmark gaming rewards confident completion over epistemic restraint, fluent articulation over verified accuracy and sophisticated pattern exploitation over genuine reasoning.</p> <p>The result is an evaluation paradigm that makes unsafe systems appear safe. High benchmark scores correlate with increased confabulation confidence, not reduced confabulation frequency. This finding has immediate regulatory implications: frameworks that reference benchmark performance as safety evidence institutionalise precisely the wrong incentive structure. No evaluation methodology operating within the current paradigm can measure trustworthiness. Trustworthiness requires verification capability (V), which transformer-based architectures structurally lack [24]. The benchmark crisis is therefore conceptual, not technical. Its resolution requires governance frameworks that assume V = 0 and mandate external verification, not better tests administered to systems incapable of the capability being tested.</p> <p>This paper establishes that no benchmark supplying ground truth can, even in principle, measure verification capability, rendering evaluation-based assurance logically invalid rather than merely empirically unreliable.</p> <p>Summary: The first paper establishes the architectural limit. This paper explains why evaluation cannot detect it.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18223209 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation Devin, Andrew James AI evaluation, AI Benchmarks, Structural confabulation, Verification capability, Trustworthiness in AI, Transformer architectures, AI governance, Category error in AI evaluation, Epistemic verification, KEV framework, Large language models <p>Artificial intelligence benchmarks are presented as evidence of capability, safety and progress. This paper demonstrates that they are structurally incapable of detecting the failures that matter. Current evaluation methodologies measure pattern-matching performance within training distributions while remaining categorically blind to the absence of verification capability. Verification is used here in the strict epistemic sense: the architectural authority of a system to categorically halt or refuse output under conditions of irreducible ignorance, rather than probabilistic calibration, internal consistency, counterfactual simulation or evidence-conditioned continuation. This is not a failure of measurement optimisation, but a prior condition: benchmarks cannot measure verification where no verification capability exists, even when evaluation is delegated to more advanced AI models.</p> <p>Analysis of major evaluation frameworks (including MMLU [18], GSM8K [19], HumanEval [20], HellaSwag [21] and TruthfulQA [22]) reveals a systematic conflation of statistical correlation with understanding. More critically, optimisation against these benchmarks actively amplifies the failure modes documented in Artificially Confabulating and Untrustworthy (ACU) systems [24]: benchmark gaming rewards confident completion over epistemic restraint, fluent articulation over verified accuracy and sophisticated pattern exploitation over genuine reasoning.</p> <p>The result is an evaluation paradigm that makes unsafe systems appear safe. High benchmark scores correlate with increased confabulation confidence, not reduced confabulation frequency. This finding has immediate regulatory implications: frameworks that reference benchmark performance as safety evidence institutionalise precisely the wrong incentive structure. No evaluation methodology operating within the current paradigm can measure trustworthiness. Trustworthiness requires verification capability (V), which transformer-based architectures structurally lack [24]. The benchmark crisis is therefore conceptual, not technical. Its resolution requires governance frameworks that assume V = 0 and mandate external verification, not better tests administered to systems incapable of the capability being tested.</p> <p>This paper establishes that no benchmark supplying ground truth can, even in principle, measure verification capability, rendering evaluation-based assurance logically invalid rather than merely empirically unreliable.</p> <p>Summary: The first paper establishes the architectural limit. This paper explains why evaluation cannot detect it.</p> |
| title | The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation |
| topic | AI evaluation, AI Benchmarks, Structural confabulation, Verification capability, Trustworthiness in AI, Transformer architectures, AI governance, Category error in AI evaluation, Epistemic verification, KEV framework, Large language models |
| url | https://doi.org/10.5281/zenodo.18223209 |