ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lior, Gili, Habba, Eliya, Levy, Shahar, Caciularu, Avi, Stanovsky, Gabriel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908536543379456
author Lior, Gili
Habba, Eliya
Levy, Shahar
Caciularu, Avi
Stanovsky, Gabriel
author_facet Lior, Gili
Habba, Eliya
Levy, Shahar
Caciularu, Avi
Stanovsky, Gabriel
contents LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt perturbations. We introduce a formal definition of reliable evaluation that accounts for prompt sensitivity, and suggest ReliableEval - a method for estimating the number of prompt resamplings needed to obtain meaningful results. Using our framework, we stochastically evaluate five frontier LLMs and find that even top-performing models like GPT-4o and Claude-3.7-Sonnet exhibit substantial prompt sensitivity. Our approach is model-, task-, and metric-agnostic, offering a recipe for meaningful and robust LLM evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22169
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
Lior, Gili
Habba, Eliya
Levy, Shahar
Caciularu, Avi
Stanovsky, Gabriel
Computation and Language
LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt perturbations. We introduce a formal definition of reliable evaluation that accounts for prompt sensitivity, and suggest ReliableEval - a method for estimating the number of prompt resamplings needed to obtain meaningful results. Using our framework, we stochastically evaluate five frontier LLMs and find that even top-performing models like GPT-4o and Claude-3.7-Sonnet exhibit substantial prompt sensitivity. Our approach is model-, task-, and metric-agnostic, offering a recipe for meaningful and robust LLM evaluation.
title ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
topic Computation and Language
url https://arxiv.org/abs/2505.22169