Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tolochinsky, Elad, Tenzer, Yaniv, Romano, Yaniv
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918494473289728
author Tolochinsky, Elad
Tenzer, Yaniv
Romano, Yaniv
author_facet Tolochinsky, Elad
Tenzer, Yaniv
Romano, Yaniv
contents Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example pair to evaluate, thereby avoiding wasted evaluations on clearly underperforming models. Further savings can be achieved by predicting model scores from the partially observed model-example score matrix using low-rank factorization. However, such predictions are not ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap predicted scores without compromising statistical validity. Specifically, we derive doubly robust estimators of each model's performance that use the low-rank predictions to reduce variance. This enables the construction of valid finite-sample confidence intervals in our setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks show that our approach reduces the number of required evaluations, yielding meaningful savings in compute and cost while accurately identifying the best-performing model.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10405
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
Tolochinsky, Elad
Tenzer, Yaniv
Romano, Yaniv
Machine Learning
Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example pair to evaluate, thereby avoiding wasted evaluations on clearly underperforming models. Further savings can be achieved by predicting model scores from the partially observed model-example score matrix using low-rank factorization. However, such predictions are not ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap predicted scores without compromising statistical validity. Specifically, we derive doubly robust estimators of each model's performance that use the low-rank predictions to reduce variance. This enables the construction of valid finite-sample confidence intervals in our setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks show that our approach reduces the number of required evaluations, yielding meaningful savings in compute and cost while accurately identifying the best-performing model.
title Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
topic Machine Learning
url https://arxiv.org/abs/2605.10405