How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Singh, Prabhant, Hess, Sibylle, Vanschoren, Joaquin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908580709400576
author Singh, Prabhant
Hess, Sibylle
Vanschoren, Joaquin
author_facet Singh, Prabhant
Hess, Sibylle
Vanschoren, Joaquin
contents Transferability estimation metrics are used to find a high-performing pre-trained model for a given target task without fine-tuning models and without access to the source dataset. Despite the growing interest in developing such metrics, the benchmarks used to measure their progress have gone largely unexamined. In this work, we empirically show the shortcomings of widely used benchmark setups to evaluate transferability estimation metrics. We argue that the benchmarks on which these metrics are evaluated are fundamentally flawed. We empirically demonstrate that their unrealistic model spaces and static performance hierarchies artificially inflate the perceived performance of existing metrics, to the point where simple, dataset-agnostic heuristics can outperform sophisticated methods. Our analysis reveals a critical disconnect between current evaluation protocols and the complexities of real-world model selection. To address this, we provide concrete recommendations for constructing more robust and realistic benchmarks to guide future research in a more meaningful direction.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
Singh, Prabhant
Hess, Sibylle
Vanschoren, Joaquin
Machine Learning
Artificial Intelligence
Transferability estimation metrics are used to find a high-performing pre-trained model for a given target task without fine-tuning models and without access to the source dataset. Despite the growing interest in developing such metrics, the benchmarks used to measure their progress have gone largely unexamined. In this work, we empirically show the shortcomings of widely used benchmark setups to evaluate transferability estimation metrics. We argue that the benchmarks on which these metrics are evaluated are fundamentally flawed. We empirically demonstrate that their unrealistic model spaces and static performance hierarchies artificially inflate the perceived performance of existing metrics, to the point where simple, dataset-agnostic heuristics can outperform sophisticated methods. Our analysis reveals a critical disconnect between current evaluation protocols and the complexities of real-world model selection. To address this, we provide concrete recommendations for constructing more robust and realistic benchmarks to guide future research in a more meaningful direction.
title How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.06448