Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jain, Suryaansh, Ahmed, Umair Z., Sahai, Shubham, Leong, Ben
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914218249289728
author Jain, Suryaansh
Ahmed, Umair Z.
Sahai, Shubham
Leong, Ben
author_facet Jain, Suryaansh
Ahmed, Umair Z.
Sahai, Shubham
Leong, Ben
contents New Large Language Models (LLMs) become available every few weeks, and modern application developers confronted with the unenviable task of having to decide if they should switch to a new model. While human evaluation remains the gold standard, it is costly and unscalable. The state-of-the-art approach is to use LLMs as evaluators ( LLM-as-a-judge), but this suffers from a critical flaw: LLMs exhibit a strong positive bias. We provide empirical evidence showing that while LLMs can identify valid outputs with high accuracy (i.e., True Positive Rate 96%), they are remarkably poor at identifying invalid ones (i.e., True Negative Rate <25%). This systematic bias, coupled with class imbalance, often leads to inflated reliability scores. While ensemble-based methods like majority voting can help, we show that they are not good enough. We introduce an optimal minority-veto strategy that is resilient to missing data and mitigates this bias to a large extent. For scenarios requiring even higher precision, we propose a novel regression-based framework that directly models the validator bias using a small set of human-annotated ground truth data. On a challenging code feedback task over 366 high-school Python programs, our regression approach reduces the maximum absolute error to just 1.2%, achieving a 2x improvement over the best-performing ensemble of 14 state-of-the-art LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
Jain, Suryaansh
Ahmed, Umair Z.
Sahai, Shubham
Leong, Ben
Artificial Intelligence
New Large Language Models (LLMs) become available every few weeks, and modern application developers confronted with the unenviable task of having to decide if they should switch to a new model. While human evaluation remains the gold standard, it is costly and unscalable. The state-of-the-art approach is to use LLMs as evaluators ( LLM-as-a-judge), but this suffers from a critical flaw: LLMs exhibit a strong positive bias. We provide empirical evidence showing that while LLMs can identify valid outputs with high accuracy (i.e., True Positive Rate 96%), they are remarkably poor at identifying invalid ones (i.e., True Negative Rate <25%). This systematic bias, coupled with class imbalance, often leads to inflated reliability scores. While ensemble-based methods like majority voting can help, we show that they are not good enough. We introduce an optimal minority-veto strategy that is resilient to missing data and mitigates this bias to a large extent. For scenarios requiring even higher precision, we propose a novel regression-based framework that directly models the validator bias using a small set of human-annotated ground truth data. On a challenging code feedback task over 366 high-school Python programs, our regression approach reduces the maximum absolute error to just 1.2%, achieving a 2x improvement over the best-performing ensemble of 14 state-of-the-art LLMs.
title Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
topic Artificial Intelligence
url https://arxiv.org/abs/2510.11822