| _version_ | 1866901074096422912 |
|---|---|
| author | Sanchez, Bryan |
| author_facet | Sanchez, Bryan |
| contents | We identify and quantify four independent mechanisms that determine success or failure of language models on multiple-choice factual benchmarks: Length Ratio (structural), Distractor Semantic Coherence (semantic), Scoring Method (measurement), and Anti-Fluency Reformulation (intervention). These four mechanisms are independent, orthogonal, and combinatorial. Together they explain approximately 95%+ of oracle variance across 69 evaluation domains. Applying all four fixes simultaneously increases baseline accuracy from 0% to 75%+ without model retraining. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19058109 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance Sanchez, Bryan LLM evaluation oracle difficulty benchmark bias measurement artifacts multiple-choice fact verification We identify and quantify four independent mechanisms that determine success or failure of language models on multiple-choice factual benchmarks: Length Ratio (structural), Distractor Semantic Coherence (semantic), Scoring Method (measurement), and Anti-Fluency Reformulation (intervention). These four mechanisms are independent, orthogonal, and combinatorial. Together they explain approximately 95%+ of oracle variance across 69 evaluation domains. Applying all four fixes simultaneously increases baseline accuracy from 0% to 75%+ without model retraining. |
| title | Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance |
| topic | LLM evaluation oracle difficulty benchmark bias measurement artifacts multiple-choice fact verification |
| url | https://doi.org/10.5281/zenodo.19058109 |