ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ashury-Tahan, Shir, Mai, Yifan, Bandel, Elron, Shmueli-Scheuer, Michal, Choshen, Leshem |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
by: Ashury-Tahan, Shir, et al.
Published: (2025)
by: Ashury-Tahan, Shir, et al.
Published: (2025)
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023)
by: Perlitz, Yotam, et al.
Published: (2023)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
Label-Efficient Model Selection for Text Generation
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
by: Bandel, Elron, et al.
Published: (2024)
by: Bandel, Elron, et al.
Published: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)
by: Kour, George, et al.
Published: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
General Agent Evaluation
by: Bandel, Elron, et al.
Published: (2026)
by: Bandel, Elron, et al.
Published: (2026)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
by: Gera, Ariel, et al.
Published: (2026)
by: Gera, Ariel, et al.
Published: (2026)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Data-driven Coreference-based Ontology Building
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Generative Large Language Models Trained for Detecting Errors in Radiology Reports
by: Sun, Cong, et al.
Published: (2025)
by: Sun, Cong, et al.
Published: (2025)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
by: Ramesh, Pratik, et al.
Published: (2026)
by: Ramesh, Pratik, et al.
Published: (2026)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025)
by: Kim, Elliot, et al.
Published: (2025)
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026)
by: Waldis, Andreas, et al.
Published: (2026)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Estimating the Error of Large Language Models at Pairwise Text Comparison
by: Li, Tianyi
Published: (2025)
by: Li, Tianyi
Published: (2025)
Concurrent Linguistic Error Detection (CLED): a New Methodology for Error Detection in Large Language Models
by: Zhu, Jinhua, et al.
Published: (2024)
by: Zhu, Jinhua, et al.
Published: (2024)
ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
by: Zeng, Jiayi, et al.
Published: (2025)
by: Zeng, Jiayi, et al.
Published: (2025)
ERC-SVD: Error-Controlled SVD for Large Language Model Compression
by: Bai, Haolei, et al.
Published: (2025)
by: Bai, Haolei, et al.
Published: (2025)
Probing for Arithmetic Errors in Language Models
by: Sun, Yucheng, et al.
Published: (2025)
by: Sun, Yucheng, et al.
Published: (2025)
CRISP: Complex Reasoning with Interpretable Step-based Plans
by: Vetzler, Matan, et al.
Published: (2025)
by: Vetzler, Matan, et al.
Published: (2025)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors
by: Yadav, Vikas, et al.
Published: (2024)
by: Yadav, Vikas, et al.
Published: (2024)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
by: Zou, Minghui, et al.
Published: (2024)
by: Zou, Minghui, et al.
Published: (2024)
Boosting of Thoughts: Trial-and-Error Problem Solving with Large Language Models
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
An In-depth Evaluation of Large Language Models in Sentence Simplification with Error-based Human Assessment
by: Wu, Xuanxin, et al.
Published: (2024)
by: Wu, Xuanxin, et al.
Published: (2024)
Similar Items
-
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026) -
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
by: Ashury-Tahan, Shir, et al.
Published: (2025) -
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023) -
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025) -
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)