TabArena: A Living Benchmark for Machine Learning on Tabular Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Erickson, Nick, Purucker, Lennart, Tschalzev, Andrej, Holzmüller, David, Desai, Prateek Mutalik, Salinas, David, Hutter, Frank
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914131909541888
author Erickson, Nick
Purucker, Lennart
Tschalzev, Andrej
Holzmüller, David
Desai, Prateek Mutalik
Salinas, David
Hutter, Frank
author_facet Erickson, Nick
Purucker, Lennart
Tschalzev, Andrej
Holzmüller, David
Desai, Prateek Mutalik
Salinas, David
Hutter, Frank
contents With the growing popularity of deep learning and foundation models for tabular data, the need for standardized and reliable benchmarks is higher than ever. However, current benchmarks are static. Their design is not updated even if flaws are discovered, model versions are updated, or new models are released. To address this, we introduce TabArena, the first continuously maintained living tabular benchmarking system. To launch TabArena, we manually curate a representative collection of datasets and well-implemented models, conduct a large-scale benchmarking study to initialize a public leaderboard, and assemble a team of experienced maintainers. Our results highlight the influence of validation method and ensembling of hyperparameter configurations to benchmark models at their full potential. While gradient-boosted trees are still strong contenders on practical tabular datasets, we observe that deep learning methods have caught up under larger time budgets with ensembling. At the same time, foundation models excel on smaller datasets. Finally, we show that ensembles across models advance the state-of-the-art in tabular machine learning. We observe that some deep learning models are overrepresented in cross-model ensembles due to validation set overfitting, and we encourage model developers to address this issue. We launch TabArena with a public leaderboard, reproducible code, and maintenance protocols to create a living benchmark available at https://tabarena.ai.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16791
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TabArena: A Living Benchmark for Machine Learning on Tabular Data
Erickson, Nick
Purucker, Lennart
Tschalzev, Andrej
Holzmüller, David
Desai, Prateek Mutalik
Salinas, David
Hutter, Frank
Machine Learning
Artificial Intelligence
With the growing popularity of deep learning and foundation models for tabular data, the need for standardized and reliable benchmarks is higher than ever. However, current benchmarks are static. Their design is not updated even if flaws are discovered, model versions are updated, or new models are released. To address this, we introduce TabArena, the first continuously maintained living tabular benchmarking system. To launch TabArena, we manually curate a representative collection of datasets and well-implemented models, conduct a large-scale benchmarking study to initialize a public leaderboard, and assemble a team of experienced maintainers. Our results highlight the influence of validation method and ensembling of hyperparameter configurations to benchmark models at their full potential. While gradient-boosted trees are still strong contenders on practical tabular datasets, we observe that deep learning methods have caught up under larger time budgets with ensembling. At the same time, foundation models excel on smaller datasets. Finally, we show that ensembles across models advance the state-of-the-art in tabular machine learning. We observe that some deep learning models are overrepresented in cross-model ensembles due to validation set overfitting, and we encourage model developers to address this issue. We launch TabArena with a public leaderboard, reproducible code, and maintenance protocols to create a living benchmark available at https://tabarena.ai.
title TabArena: A Living Benchmark for Machine Learning on Tabular Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.16791