Submodular Benchmark Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Smola, Alexander
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915976821342208
author Smola, Alexander
author_facet Smola, Alexander
contents Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small, informative subset as submodular maximization under a multivariate Gaussian model. Entropy (log-determinant covariance) and mutual information between selected and remaining benchmarks arise as natural objectives. Both are submodular; entropy selection coincides with pivoted Cholesky and has spectral residual bounds, while mutual information is non-monotone in general but empirically monotone for small subsets, so we optimize it greedily. Experiments on three matrices from ten public leaderboards show that mutual information selection outperforms entropy for imputation at small subsets.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02209
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Submodular Benchmark Selection
Smola, Alexander
Artificial Intelligence
Machine Learning
90C27 (Primary), 62K05, 94A17, 68T05 (Secondary)
I.2.6; G.3; F.2.2
Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small, informative subset as submodular maximization under a multivariate Gaussian model. Entropy (log-determinant covariance) and mutual information between selected and remaining benchmarks arise as natural objectives. Both are submodular; entropy selection coincides with pivoted Cholesky and has spectral residual bounds, while mutual information is non-monotone in general but empirically monotone for small subsets, so we optimize it greedily. Experiments on three matrices from ten public leaderboards show that mutual information selection outperforms entropy for imputation at small subsets.
title Submodular Benchmark Selection
topic Artificial Intelligence
Machine Learning
90C27 (Primary), 62K05, 94A17, 68T05 (Secondary)
I.2.6; G.3; F.2.2
url https://arxiv.org/abs/2605.02209