Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Acharyya, Aranyak, Priebe, Carey E., Helm, Hayden S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911153012080640
author Acharyya, Aranyak
Priebe, Carey E.
Helm, Hayden S.
author_facet Acharyya, Aranyak
Priebe, Carey E.
Helm, Hayden S.
contents Given an input query, generative models such as large language models produce a random response drawn from a response distribution. Given two input queries, it is natural to ask if their response distributions are the same. While traditional statistical hypothesis testing is designed to address this question, the response distribution induced by an input query is often sensitive to semantically irrelevant perturbations to the query, so much so that a traditional test of equality might indicate that two semantically equivalent queries induce statistically different response distributions. As a result, the outcome of the statistical test may not align with the user's requirements. In this paper, we address this misalignment by incorporating into the testing procedure consideration of a collection of semantically similar queries. In our setting, the mapping from the collection of user-defined semantically similar queries to the corresponding collection of response distributions is not known a priori and must be estimated, with a fixed budget. Although the problem we address is quite general, we focus our analysis on the setting where the responses are binary, show that the proposed test is asymptotically valid and consistent, and discuss important practical considerations with respect to power and computation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations
Acharyya, Aranyak
Priebe, Carey E.
Helm, Hayden S.
Statistics Theory
Artificial Intelligence
Methodology
Given an input query, generative models such as large language models produce a random response drawn from a response distribution. Given two input queries, it is natural to ask if their response distributions are the same. While traditional statistical hypothesis testing is designed to address this question, the response distribution induced by an input query is often sensitive to semantically irrelevant perturbations to the query, so much so that a traditional test of equality might indicate that two semantically equivalent queries induce statistically different response distributions. As a result, the outcome of the statistical test may not align with the user's requirements. In this paper, we address this misalignment by incorporating into the testing procedure consideration of a collection of semantically similar queries. In our setting, the mapping from the collection of user-defined semantically similar queries to the corresponding collection of response distributions is not known a priori and must be estimated, with a fixed budget. Although the problem we address is quite general, we focus our analysis on the setting where the responses are binary, show that the proposed test is asymptotically valid and consistent, and discuss important practical considerations with respect to power and computation.
title Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations
topic Statistics Theory
Artificial Intelligence
Methodology
url https://arxiv.org/abs/2509.10963