Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heineman, David, Hofmann, Valentin, Magnusson, Ian, Gu, Yuling, Smith, Noah A., Hajishirzi, Hannaneh, Lo, Kyle, Dodge, Jesse
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909741062553600
author Heineman, David
Hofmann, Valentin
Magnusson, Ian
Gu, Yuling
Smith, Noah A.
Hajishirzi, Hannaneh
Lo, Kyle
Dodge, Jesse
author_facet Heineman, David
Hofmann, Valentin
Magnusson, Ian
Gu, Yuling
Smith, Noah A.
Hajishirzi, Hannaneh
Lo, Kyle
Dodge, Jesse
contents Developing large language models is expensive and involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more reliable for such decisions, and interventions to design higher-quality evaluation benchmarks. We introduce two key metrics that show differences in current benchmarks: signal, a benchmark's ability to separate better models from worse models, and noise, a benchmark's sensitivity to random variability between training steps. We demonstrate that benchmarks with a better signal-to-noise ratio are more reliable when making decisions at small scale, and those with less noise have lower scaling law prediction error. These results suggest that improving signal or noise will lead to more useful benchmarks, so we introduce three interventions designed to directly affect signal or noise. For example, we propose that switching to a metric that has better signal and noise (e.g., perplexity rather than accuracy) leads to better reliability and improved scaling law error. We also find that filtering noisy subtasks, to improve an aggregate signal-to-noise ratio, leads to more reliable multi-task evaluations. We also find that averaging the output of a model's intermediate checkpoints to reduce noise leads to consistent improvements. We conclude by recommending that those creating new benchmarks, or selecting which existing benchmarks to use, aim for high signal and low noise. We use 30 benchmarks for these experiments, and 375 open-weight language models from 60M to 32B parameters, resulting in a new, publicly available dataset of 900K evaluation benchmark results, totaling 200M instances.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
Heineman, David
Hofmann, Valentin
Magnusson, Ian
Gu, Yuling
Smith, Noah A.
Hajishirzi, Hannaneh
Lo, Kyle
Dodge, Jesse
Computation and Language
Machine Learning
Developing large language models is expensive and involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more reliable for such decisions, and interventions to design higher-quality evaluation benchmarks. We introduce two key metrics that show differences in current benchmarks: signal, a benchmark's ability to separate better models from worse models, and noise, a benchmark's sensitivity to random variability between training steps. We demonstrate that benchmarks with a better signal-to-noise ratio are more reliable when making decisions at small scale, and those with less noise have lower scaling law prediction error. These results suggest that improving signal or noise will lead to more useful benchmarks, so we introduce three interventions designed to directly affect signal or noise. For example, we propose that switching to a metric that has better signal and noise (e.g., perplexity rather than accuracy) leads to better reliability and improved scaling law error. We also find that filtering noisy subtasks, to improve an aggregate signal-to-noise ratio, leads to more reliable multi-task evaluations. We also find that averaging the output of a model's intermediate checkpoints to reduce noise leads to consistent improvements. We conclude by recommending that those creating new benchmarks, or selecting which existing benchmarks to use, aim for high signal and low noise. We use 30 benchmarks for these experiments, and 375 open-weight language models from 60M to 32B parameters, resulting in a new, publicly available dataset of 900K evaluation benchmark results, totaling 200M instances.
title Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.13144