An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yamauchi, Yusuke, Yano, Taro, Oyamada, Masafumi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911007530549248
author Yamauchi, Yusuke
Yano, Taro
Oyamada, Masafumi
author_facet Yamauchi, Yusuke
Yano, Taro
Oyamada, Masafumi
contents As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its reliability remains uncertain. In this work, we analyze key factors affecting its trustworthiness, focusing on alignment with human judgments and evaluation consistency. Using BIGGENBench and EvalBiasBench, we study the effects of evaluation design, decoding strategies, and Chain-of-Tought (CoT) reasoning in evaluation. Our results show that evaluation criteria are critical for reliability, non-deterministic sampling improves alignment with human preferences over deterministic evaluation, and CoT reasoning offers minimal gains when clear evaluation criteria are present.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Yamauchi, Yusuke
Yano, Taro
Oyamada, Masafumi
Computation and Language
As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its reliability remains uncertain. In this work, we analyze key factors affecting its trustworthiness, focusing on alignment with human judgments and evaluation consistency. Using BIGGENBench and EvalBiasBench, we study the effects of evaluation design, decoding strategies, and Chain-of-Tought (CoT) reasoning in evaluation. Our results show that evaluation criteria are critical for reliability, non-deterministic sampling improves alignment with human preferences over deterministic evaluation, and CoT reasoning offers minimal gains when clear evaluation criteria are present.
title An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
topic Computation and Language
url https://arxiv.org/abs/2506.13639