Bayesian Evaluation of Large Language Model Behavior

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Longjohn, Rachel, Wu, Shang, Kher, Saatvik, Belém, Catarina, Smyth, Padhraic
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911264148553728
author Longjohn, Rachel
Wu, Shang
Kher, Saatvik
Belém, Catarina
Smyth, Padhraic
author_facet Longjohn, Rachel
Wu, Shang
Kher, Saatvik
Belém, Catarina
Smyth, Padhraic
contents It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensitivity to adversarial inputs. Such evaluations often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assessed in a binary fashion (e.g., harmful/non-harmful or does not leak/leaks sensitive information), and the aggregation of binary scores is used to evaluate the LLM. However, existing approaches to evaluation often neglect statistical uncertainty quantification. With an applied statistics audience in mind, we provide background on LLM text generation and evaluation, and then describe a Bayesian approach for quantifying uncertainty in binary evaluation metrics. We focus in particular on uncertainty that is induced by the probabilistic text generation strategies typically deployed in LLM-based systems. We present two case studies applying this approach: 1) evaluating refusal rates on a benchmark of adversarial inputs designed to elicit harmful responses, and 2) evaluating pairwise preferences of one LLM over another on a benchmark of open-ended interactive dialogue examples. We demonstrate how the Bayesian approach can provide useful uncertainty quantification about the behavior of LLM-based systems.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bayesian Evaluation of Large Language Model Behavior
Longjohn, Rachel
Wu, Shang
Kher, Saatvik
Belém, Catarina
Smyth, Padhraic
Computation and Language
Machine Learning
Applications
It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensitivity to adversarial inputs. Such evaluations often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assessed in a binary fashion (e.g., harmful/non-harmful or does not leak/leaks sensitive information), and the aggregation of binary scores is used to evaluate the LLM. However, existing approaches to evaluation often neglect statistical uncertainty quantification. With an applied statistics audience in mind, we provide background on LLM text generation and evaluation, and then describe a Bayesian approach for quantifying uncertainty in binary evaluation metrics. We focus in particular on uncertainty that is induced by the probabilistic text generation strategies typically deployed in LLM-based systems. We present two case studies applying this approach: 1) evaluating refusal rates on a benchmark of adversarial inputs designed to elicit harmful responses, and 2) evaluating pairwise preferences of one LLM over another on a benchmark of open-ended interactive dialogue examples. We demonstrate how the Bayesian approach can provide useful uncertainty quantification about the behavior of LLM-based systems.
title Bayesian Evaluation of Large Language Model Behavior
topic Computation and Language
Machine Learning
Applications
url https://arxiv.org/abs/2511.10661