CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Ziyi, Tieleman, Olivier, Bukhtiyarov, Alexey, Chen, Jinghong
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913086593564672
author Zhu, Ziyi
Tieleman, Olivier
Bukhtiyarov, Alexey
Chen, Jinghong
author_facet Zhu, Ziyi
Tieleman, Olivier
Bukhtiyarov, Alexey
Chen, Jinghong
contents LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number of scenarios or generations. These biases are often similar in magnitude to the model differences that benchmarks are designed to detect, resulting in unreliable rankings when single-judge evaluations are used. We introduce a variance decomposition that partitions benchmark score variance into scenario, generation, judge, and residual components. Based on this analysis, CyclicJudge, a round-robin assignment of judges to scenarios, is demonstrated to be the optimal strategy for a fixed judge panel and judge-call budget: the score recovers the panel mean exactly while matching the cost of single-judge evaluation. Empirical results on MT-Bench and MindEval validate the effectiveness of CyclicJudge as predicted, across both general-purpose and domain-specific evaluation settings.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01865
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
Zhu, Ziyi
Tieleman, Olivier
Bukhtiyarov, Alexey
Chen, Jinghong
Computation and Language
LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number of scenarios or generations. These biases are often similar in magnitude to the model differences that benchmarks are designed to detect, resulting in unreliable rankings when single-judge evaluations are used. We introduce a variance decomposition that partitions benchmark score variance into scenario, generation, judge, and residual components. Based on this analysis, CyclicJudge, a round-robin assignment of judges to scenarios, is demonstrated to be the optimal strategy for a fixed judge panel and judge-call budget: the score recovers the panel mean exactly while matching the cost of single-judge evaluation. Empirical results on MT-Bench and MindEval validate the effectiveness of CyclicJudge as predicted, across both general-purpose and domain-specific evaluation settings.
title CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
topic Computation and Language
url https://arxiv.org/abs/2603.01865