Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Wenbo, Cai, Hengrui, Chen, Wenyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915999273451520
author Zhang, Wenbo
Cai, Hengrui
Chen, Wenyu
author_facet Zhang, Wenbo
Cai, Hengrui
Chen, Wenyu
contents Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural language processing and understanding. Benchmark evaluations are crucial for assessing the capabilities of LLMs as they can provide a comprehensive assessment of their strengths and weaknesses. However, current evaluation methods often overlook the inherent randomness of LLMs by employing deterministic generation strategies or relying on a single random sample, resulting in unaccounted sampling variance and unreliable benchmark score estimates. In this paper, we propose a hierarchical statistical model that provides a more comprehensive representation of the benchmarking process by incorporating both benchmark characteristics and LLM randomness. We show that leveraging multiple generations improves the accuracy of estimating the benchmark score and reduces variance. Multiple generations also allow us to define $\mathbb P\left(\text{correct}\right)$, a prompt-level difficulty score based on correct ratios, providing fine-grained insights into individual prompts. Additionally, we create a data map that visualizes difficulty and semantics of prompts, enabling error detection and quality control in benchmark construction.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
Zhang, Wenbo
Cai, Hengrui
Chen, Wenyu
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural language processing and understanding. Benchmark evaluations are crucial for assessing the capabilities of LLMs as they can provide a comprehensive assessment of their strengths and weaknesses. However, current evaluation methods often overlook the inherent randomness of LLMs by employing deterministic generation strategies or relying on a single random sample, resulting in unaccounted sampling variance and unreliable benchmark score estimates. In this paper, we propose a hierarchical statistical model that provides a more comprehensive representation of the benchmarking process by incorporating both benchmark characteristics and LLM randomness. We show that leveraging multiple generations improves the accuracy of estimating the benchmark score and reduces variance. Multiple generations also allow us to define $\mathbb P\left(\text{correct}\right)$, a prompt-level difficulty score based on correct ratios, providing fine-grained insights into individual prompts. Additionally, we create a data map that visualizes difficulty and semantics of prompts, enabling error detection and quality control in benchmark construction.
title Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.08943