An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911007530549248 |
|---|---|
| author | Yamauchi, Yusuke Yano, Taro Oyamada, Masafumi |
| author_facet | Yamauchi, Yusuke Yano, Taro Oyamada, Masafumi |
| contents | As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its reliability remains uncertain. In this work, we analyze key factors affecting its trustworthiness, focusing on alignment with human judgments and evaluation consistency. Using BIGGENBench and EvalBiasBench, we study the effects of evaluation design, decoding strategies, and Chain-of-Tought (CoT) reasoning in evaluation. Our results show that evaluation criteria are critical for reliability, non-deterministic sampling improves alignment with human preferences over deterministic evaluation, and CoT reasoning offers minimal gains when clear evaluation criteria are present. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_13639 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Yamauchi, Yusuke Yano, Taro Oyamada, Masafumi Computation and Language As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its reliability remains uncertain. In this work, we analyze key factors affecting its trustworthiness, focusing on alignment with human judgments and evaluation consistency. Using BIGGENBench and EvalBiasBench, we study the effects of evaluation design, decoding strategies, and Chain-of-Tought (CoT) reasoning in evaluation. Our results show that evaluation criteria are critical for reliability, non-deterministic sampling improves alignment with human preferences over deterministic evaluation, and CoT reasoning offers minimal gains when clear evaluation criteria are present. |
| title | An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2506.13639 |