CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yehudai, Asaf, Eden, Lilach, Perlitz, Yotam, Bar-Haim, Roy, Shmueli-Scheuer, Michal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909703711227904
author Yehudai, Asaf
Eden, Lilach
Perlitz, Yotam
Bar-Haim, Roy
Shmueli-Scheuer, Michal
author_facet Yehudai, Asaf
Eden, Lilach
Perlitz, Yotam
Bar-Haim, Roy
Shmueli-Scheuer, Michal
contents The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model's performance. To bridge this gap, we introduce CLEAR, an interactive, open-source package for LLM-based error analysis. CLEAR first generates per-instance textual feedback, then it creates a set of system-level error issues, and quantifies the prevalence of each identified issue. Our package also provides users with an interactive dashboard that allows for a comprehensive error analysis through aggregate visualizations, applies interactive filters to isolate specific issues or score ranges, and drills down to the individual instances that exemplify a particular behavioral pattern. We demonstrate CLEAR analysis for RAG and Math benchmarks, and showcase its utility through a user case study.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
Yehudai, Asaf
Eden, Lilach
Perlitz, Yotam
Bar-Haim, Roy
Shmueli-Scheuer, Michal
Computation and Language
Artificial Intelligence
Machine Learning
The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model's performance. To bridge this gap, we introduce CLEAR, an interactive, open-source package for LLM-based error analysis. CLEAR first generates per-instance textual feedback, then it creates a set of system-level error issues, and quantifies the prevalence of each identified issue. Our package also provides users with an interactive dashboard that allows for a comprehensive error analysis through aggregate visualizations, applies interactive filters to isolate specific issues or score ranges, and drills down to the individual instances that exemplify a particular behavioral pattern. We demonstrate CLEAR analysis for RAG and Math benchmarks, and showcase its utility through a user case study.
title CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.18392