carps: A Framework for Comparing N Hyperparameter Optimizers on M Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benjamins, Carolin, Graf, Helena, Segel, Sarah, Deng, Difan, Ruhkopf, Tim, Hennig, Leona, Basu, Soham, Mallik, Neeratyoy, Bergman, Edward, Chen, Deyao, Clément, François, Tornede, Alexander, Feurer, Matthias, Eggensperger, Katharina, Hutter, Frank, Doerr, Carola, Lindauer, Marius
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918143216058368
author Benjamins, Carolin
Graf, Helena
Segel, Sarah
Deng, Difan
Ruhkopf, Tim
Hennig, Leona
Basu, Soham
Mallik, Neeratyoy
Bergman, Edward
Chen, Deyao
Clément, François
Tornede, Alexander
Feurer, Matthias
Eggensperger, Katharina
Hutter, Frank
Doerr, Carola
Lindauer, Marius
author_facet Benjamins, Carolin
Graf, Helena
Segel, Sarah
Deng, Difan
Ruhkopf, Tim
Hennig, Leona
Basu, Soham
Mallik, Neeratyoy
Bergman, Edward
Chen, Deyao
Clément, François
Tornede, Alexander
Feurer, Matthias
Eggensperger, Katharina
Hutter, Frank
Doerr, Carola
Lindauer, Marius
contents Hyperparameter Optimization (HPO) is crucial to develop well-performing machine learning models. In order to ease prototyping and benchmarking of HPO methods, we propose carps, a benchmark framework for Comprehensive Automated Research Performance Studies allowing to evaluate N optimizers on M benchmark tasks. In this first release of carps, we focus on the four most important types of HPO task types: blackbox, multi-fidelity, multi-objective and multi-fidelity-multi-objective. With 3 336 tasks from 5 community benchmark collections and 28 variants of 9 optimizer families, we offer the biggest go-to library to date to evaluate and compare HPO methods. The carps framework relies on a purpose-built, lightweight interface, gluing together optimizers and benchmark tasks. It also features an analysis pipeline, facilitating the evaluation of optimizers on benchmarks. However, navigating a huge number of tasks while developing and comparing methods can be computationally infeasible. To address this, we obtain a subset of representative tasks by minimizing the star discrepancy of the subset, in the space spanned by the full set. As a result, we propose an initial subset of 10 to 30 diverse tasks for each task type, and include functionality to re-compute subsets as more benchmarks become available, enabling efficient evaluations. We also establish a first set of baseline results on these tasks as a measure for future comparisons. With carps (https://www.github.com/automl/CARP-S), we make an important step in the standardization of HPO evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle carps: A Framework for Comparing N Hyperparameter Optimizers on M Benchmarks
Benjamins, Carolin
Graf, Helena
Segel, Sarah
Deng, Difan
Ruhkopf, Tim
Hennig, Leona
Basu, Soham
Mallik, Neeratyoy
Bergman, Edward
Chen, Deyao
Clément, François
Tornede, Alexander
Feurer, Matthias
Eggensperger, Katharina
Hutter, Frank
Doerr, Carola
Lindauer, Marius
Machine Learning
Hyperparameter Optimization (HPO) is crucial to develop well-performing machine learning models. In order to ease prototyping and benchmarking of HPO methods, we propose carps, a benchmark framework for Comprehensive Automated Research Performance Studies allowing to evaluate N optimizers on M benchmark tasks. In this first release of carps, we focus on the four most important types of HPO task types: blackbox, multi-fidelity, multi-objective and multi-fidelity-multi-objective. With 3 336 tasks from 5 community benchmark collections and 28 variants of 9 optimizer families, we offer the biggest go-to library to date to evaluate and compare HPO methods. The carps framework relies on a purpose-built, lightweight interface, gluing together optimizers and benchmark tasks. It also features an analysis pipeline, facilitating the evaluation of optimizers on benchmarks. However, navigating a huge number of tasks while developing and comparing methods can be computationally infeasible. To address this, we obtain a subset of representative tasks by minimizing the star discrepancy of the subset, in the space spanned by the full set. As a result, we propose an initial subset of 10 to 30 diverse tasks for each task type, and include functionality to re-compute subsets as more benchmarks become available, enabling efficient evaluations. We also establish a first set of baseline results on these tasks as a measure for future comparisons. With carps (https://www.github.com/automl/CARP-S), we make an important step in the standardization of HPO evaluation.
title carps: A Framework for Comparing N Hyperparameter Optimizers on M Benchmarks
topic Machine Learning
url https://arxiv.org/abs/2506.06143