Submodular Benchmark Selection
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915976821342208 |
|---|---|
| author | Smola, Alexander |
| author_facet | Smola, Alexander |
| contents | Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small, informative subset as submodular maximization under a multivariate Gaussian model. Entropy (log-determinant covariance) and mutual information between selected and remaining benchmarks arise as natural objectives. Both are submodular; entropy selection coincides with pivoted Cholesky and has spectral residual bounds, while mutual information is non-monotone in general but empirically monotone for small subsets, so we optimize it greedily. Experiments on three matrices from ten public leaderboards show that mutual information selection outperforms entropy for imputation at small subsets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_02209 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Submodular Benchmark Selection Smola, Alexander Artificial Intelligence Machine Learning 90C27 (Primary), 62K05, 94A17, 68T05 (Secondary) I.2.6; G.3; F.2.2 Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small, informative subset as submodular maximization under a multivariate Gaussian model. Entropy (log-determinant covariance) and mutual information between selected and remaining benchmarks arise as natural objectives. Both are submodular; entropy selection coincides with pivoted Cholesky and has spectral residual bounds, while mutual information is non-monotone in general but empirically monotone for small subsets, so we optimize it greedily. Experiments on three matrices from ten public leaderboards show that mutual information selection outperforms entropy for imputation at small subsets. |
| title | Submodular Benchmark Selection |
| topic | Artificial Intelligence Machine Learning 90C27 (Primary), 62K05, 94A17, 68T05 (Secondary) I.2.6; G.3; F.2.2 |
| url | https://arxiv.org/abs/2605.02209 |