Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Sanchez, Bryan
Format: Recurso digital
Published: Zenodo 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901074096422912
author Sanchez, Bryan
author_facet Sanchez, Bryan
contents We identify and quantify four independent mechanisms that determine success or failure of language models on multiple-choice factual benchmarks: Length Ratio (structural), Distractor Semantic Coherence (semantic), Scoring Method (measurement), and Anti-Fluency Reformulation (intervention). These four mechanisms are independent, orthogonal, and combinatorial. Together they explain approximately 95%+ of oracle variance across 69 evaluation domains. Applying all four fixes simultaneously increases baseline accuracy from 0% to 75%+ without model retraining.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19058109
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance
Sanchez, Bryan
LLM evaluation
oracle difficulty
benchmark bias
measurement artifacts
multiple-choice
fact verification
We identify and quantify four independent mechanisms that determine success or failure of language models on multiple-choice factual benchmarks: Length Ratio (structural), Distractor Semantic Coherence (semantic), Scoring Method (measurement), and Anti-Fluency Reformulation (intervention). These four mechanisms are independent, orthogonal, and combinatorial. Together they explain approximately 95%+ of oracle variance across 69 evaluation domains. Applying all four fixes simultaneously increases baseline accuracy from 0% to 75%+ without model retraining.
title Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance
topic LLM evaluation
oracle difficulty
benchmark bias
measurement artifacts
multiple-choice
fact verification
url https://doi.org/10.5281/zenodo.19058109