Estimating LLM Consistency: A User Baseline vs Surrogate Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiaoyuan, Lin, Weiran, Akgul, Omer, Bauer, Lujo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915631634317312
author Wu, Xiaoyuan
Lin, Weiran
Akgul, Omer
Bauer, Lujo
author_facet Wu, Xiaoyuan
Lin, Weiran
Akgul, Omer
Bauer, Lujo
contents Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text. Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM responses -- the model's confidence in the response or likelihood of generating a similar response when resampled. In previous work, measuring LLM response consistency often relied on calculating the probability of a response appearing within a pool of resampled responses, analyzing internal states, or evaluating logits of responses. However, it was not clear how well these approaches approximated users' perceptions of consistency of LLM responses. To find out, we performed a user study ($n=2,976$) demonstrating that current methods for measuring LLM response consistency typically do not align well with humans' perceptions of LLM consistency. We propose a logit-based ensemble method for estimating LLM consistency and show that our method matches the performance of the best-performing existing metric in estimating human ratings of LLM consistency. Our results suggest that methods for estimating LLM consistency without human evaluation are sufficiently imperfect to warrant broader use of evaluation with human input; this would avoid misjudging the adequacy of models because of the imperfections of automated consistency metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
Wu, Xiaoyuan
Lin, Weiran
Akgul, Omer
Bauer, Lujo
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text. Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM responses -- the model's confidence in the response or likelihood of generating a similar response when resampled. In previous work, measuring LLM response consistency often relied on calculating the probability of a response appearing within a pool of resampled responses, analyzing internal states, or evaluating logits of responses. However, it was not clear how well these approaches approximated users' perceptions of consistency of LLM responses. To find out, we performed a user study ($n=2,976$) demonstrating that current methods for measuring LLM response consistency typically do not align well with humans' perceptions of LLM consistency. We propose a logit-based ensemble method for estimating LLM consistency and show that our method matches the performance of the best-performing existing metric in estimating human ratings of LLM consistency. Our results suggest that methods for estimating LLM consistency without human evaluation are sufficiently imperfect to warrant broader use of evaluation with human input; this would avoid misjudging the adequacy of models because of the imperfections of automated consistency metrics.
title Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2505.23799