Saved in:
Bibliographic Details
Main Authors: Blackwell, Robert E., Barry, Jon, Cohn, Anthony G.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.03492
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908424485208064
author Blackwell, Robert E.
Barry, Jon
Cohn, Anthony G.
author_facet Blackwell, Robert E.
Barry, Jon
Cohn, Anthony G.
contents Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the time and cost of repeated experiments. We use benchmarks designed for testing LLMs' capacity to reason about cardinal directions to explore the impact of experimental repeats on mean score and prediction interval. We suggest a simple method for cost-effectively quantifying the uncertainty of a benchmark score and make recommendations concerning reproducible LLM evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03492
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
Blackwell, Robert E.
Barry, Jon
Cohn, Anthony G.
Computation and Language
Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the time and cost of repeated experiments. We use benchmarks designed for testing LLMs' capacity to reason about cardinal directions to explore the impact of experimental repeats on mean score and prediction interval. We suggest a simple method for cost-effectively quantifying the uncertainty of a benchmark score and make recommendations concerning reproducible LLM evaluation.
title Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
topic Computation and Language
url https://arxiv.org/abs/2410.03492