ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ashury-Tahan, Shir, Mai, Yifan, Bandel, Elron, Shmueli-Scheuer, Michal, Choshen, Leshem
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914333702750208
author Ashury-Tahan, Shir
Mai, Yifan
Bandel, Elron
Shmueli-Scheuer, Michal
Choshen, Leshem
author_facet Ashury-Tahan, Shir
Mai, Yifan
Bandel, Elron
Shmueli-Scheuer, Michal
Choshen, Leshem
contents Large Language Models (LLM) benchmarks tell us when models fail, but not why they fail. A wrong answer on a reasoning dataset may stem from formatting issues, calculation errors, or dataset noise rather than weak reasoning. Without disentangling such causes, benchmarks remain incomplete and cannot reliably guide model improvement. We introduce ErrorMap, the first method to chart the sources of LLM failure. It extracts a model's unique "failure signature", clarifies what benchmarks measure, and broadens error identification to reduce blind spots. This helps developers debug models, aligns benchmark goals with outcomes, and supports informed model selection. ErrorMap works on any model or dataset with the same logic. Applying our method to 35 datasets and 83 models we generate ErrorAtlas, a taxonomy of model errors, revealing recurring failure patterns. ErrorAtlas highlights error types that are currently underexplored in LLM research, such as omissions of required details in the output and question misinterpretation. By shifting focus from where models succeed to why they fail, ErrorMap and ErrorAtlas enable advanced evaluation - one that exposes hidden weaknesses and directs progress. Unlike success, typically measured by task-level metrics, our approach introduces a deeper evaluation layer that can be applied globally across models and tasks, offering richer insights into model behavior and limitations. We make the taxonomy and code publicly available with plans to periodically update ErrorAtlas as new benchmarks and models emerge.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15812
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
Ashury-Tahan, Shir
Mai, Yifan
Bandel, Elron
Shmueli-Scheuer, Michal
Choshen, Leshem
Artificial Intelligence
Computation and Language
Large Language Models (LLM) benchmarks tell us when models fail, but not why they fail. A wrong answer on a reasoning dataset may stem from formatting issues, calculation errors, or dataset noise rather than weak reasoning. Without disentangling such causes, benchmarks remain incomplete and cannot reliably guide model improvement. We introduce ErrorMap, the first method to chart the sources of LLM failure. It extracts a model's unique "failure signature", clarifies what benchmarks measure, and broadens error identification to reduce blind spots. This helps developers debug models, aligns benchmark goals with outcomes, and supports informed model selection. ErrorMap works on any model or dataset with the same logic. Applying our method to 35 datasets and 83 models we generate ErrorAtlas, a taxonomy of model errors, revealing recurring failure patterns. ErrorAtlas highlights error types that are currently underexplored in LLM research, such as omissions of required details in the output and question misinterpretation. By shifting focus from where models succeed to why they fail, ErrorMap and ErrorAtlas enable advanced evaluation - one that exposes hidden weaknesses and directs progress. Unlike success, typically measured by task-level metrics, our approach introduces a deeper evaluation layer that can be applied globally across models and tasks, offering richer insights into model behavior and limitations. We make the taxonomy and code publicly available with plans to periodically update ErrorAtlas as new benchmarks and models emerge.
title ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.15812