Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haldar, Rajarshi, Hockenmaier, Julia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915961641107456
author Haldar, Rajarshi
Hockenmaier, Julia
author_facet Haldar, Rajarshi
Hockenmaier, Julia
contents As Natural Language Generation (NLG) continues to be widely adopted, properly assessing it has become quite difficult. Lately, using large language models (LLMs) for evaluating these generations has gained traction, as they tend to align more closely with human preferences than conventional n-gram or embedding-based metrics. In our experiments, we show that LLM judges have low intra-rater reliability in their assigned scores across different runs. This variance makes their ratings inconsistent, almost arbitrary in the worst case, making it difficult to measure how good their judgments actually are. We quantify this inconsistency across different NLG tasks and benchmarks and see if judicious use of LLM judges can still be useful following proper guidelines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
Haldar, Rajarshi
Hockenmaier, Julia
Computation and Language
As Natural Language Generation (NLG) continues to be widely adopted, properly assessing it has become quite difficult. Lately, using large language models (LLMs) for evaluating these generations has gained traction, as they tend to align more closely with human preferences than conventional n-gram or embedding-based metrics. In our experiments, we show that LLM judges have low intra-rater reliability in their assigned scores across different runs. This variance makes their ratings inconsistent, almost arbitrary in the worst case, making it difficult to measure how good their judgments actually are. We quantify this inconsistency across different NLG tasks and benchmarks and see if judicious use of LLM judges can still be useful following proper guidelines.
title Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
topic Computation and Language
url https://arxiv.org/abs/2510.27106