ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908536543379456 |
|---|---|
| author | Lior, Gili Habba, Eliya Levy, Shahar Caciularu, Avi Stanovsky, Gabriel |
| author_facet | Lior, Gili Habba, Eliya Levy, Shahar Caciularu, Avi Stanovsky, Gabriel |
| contents | LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt perturbations. We introduce a formal definition of reliable evaluation that accounts for prompt sensitivity, and suggest ReliableEval - a method for estimating the number of prompt resamplings needed to obtain meaningful results. Using our framework, we stochastically evaluate five frontier LLMs and find that even top-performing models like GPT-4o and Claude-3.7-Sonnet exhibit substantial prompt sensitivity. Our approach is model-, task-, and metric-agnostic, offering a recipe for meaningful and robust LLM evaluation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_22169 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments Lior, Gili Habba, Eliya Levy, Shahar Caciularu, Avi Stanovsky, Gabriel Computation and Language LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments evaluation over the space of meaning-preserving prompt perturbations. We introduce a formal definition of reliable evaluation that accounts for prompt sensitivity, and suggest ReliableEval - a method for estimating the number of prompt resamplings needed to obtain meaningful results. Using our framework, we stochastically evaluate five frontier LLMs and find that even top-performing models like GPT-4o and Claude-3.7-Sonnet exhibit substantial prompt sensitivity. Our approach is model-, task-, and metric-agnostic, offering a recipe for meaningful and robust LLM evaluation. |
| title | ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.22169 |