Improving LLM-as-a-Judge Inference with the Judgment Distribution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Victor, Zhang, Michael J. Q., Choi, Eunsol
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914056843034624
author Wang, Victor
Zhang, Michael J. Q.
Choi, Eunsol
author_facet Wang, Victor
Zhang, Michael J. Q.
Choi, Eunsol
contents Using language models to scalably approximate human preferences on text quality (LLM-as-a-judge) has become a standard practice applicable to many tasks. A judgment is often extracted from the judge's textual output alone, typically with greedy decoding. However, LLM judges naturally provide distributions over judgment tokens, inviting a breadth of inference methods for extracting fine-grained preferences. We find that taking the mean of the judgment distribution consistently outperforms taking the mode (i.e. greedy decoding) in all evaluation settings (i.e. pointwise, pairwise, and listwise). We further explore novel methods of deriving preferences from judgment distributions, and find that methods incorporating risk aversion often improve performance. Lastly, we analyze LLM-as-a-judge paired with chain-of-thought (CoT) prompting, showing that CoT can collapse the spread of the judgment distribution, often harming performance. Our findings show that leveraging distributional output improves LLM-as-a-judge, as opposed to using the text interface alone.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03064
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving LLM-as-a-Judge Inference with the Judgment Distribution
Wang, Victor
Zhang, Michael J. Q.
Choi, Eunsol
Computation and Language
Using language models to scalably approximate human preferences on text quality (LLM-as-a-judge) has become a standard practice applicable to many tasks. A judgment is often extracted from the judge's textual output alone, typically with greedy decoding. However, LLM judges naturally provide distributions over judgment tokens, inviting a breadth of inference methods for extracting fine-grained preferences. We find that taking the mean of the judgment distribution consistently outperforms taking the mode (i.e. greedy decoding) in all evaluation settings (i.e. pointwise, pairwise, and listwise). We further explore novel methods of deriving preferences from judgment distributions, and find that methods incorporating risk aversion often improve performance. Lastly, we analyze LLM-as-a-judge paired with chain-of-thought (CoT) prompting, showing that CoT can collapse the spread of the judgment distribution, often harming performance. Our findings show that leveraging distributional output improves LLM-as-a-judge, as opposed to using the text interface alone.
title Improving LLM-as-a-Judge Inference with the Judgment Distribution
topic Computation and Language
url https://arxiv.org/abs/2503.03064