Saved in:
Bibliographic Details
Main Authors: Liu, Ruizhe, Luo, Jiaqi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.14915
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918501725241344
author Liu, Ruizhe
Luo, Jiaqi
author_facet Liu, Ruizhe
Luo, Jiaqi
contents Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed algorithms, a systematic empirical understanding of how different imbalanced learning methods behave across diverse data characteristics is still lacking. In particular, it remains unclear how different method families compare in predictive performance, robustness under varying data characteristics, and computational scalability. In this work, we present Tabular Imbalanced Learning Benchmark (TILBench), a large-scale empirical benchmark for tabular imbalanced learning. TILBench evaluates more than 40 representative algorithms across 57 diverse tabular datasets, resulting in over 200000 controlled experiments across a wide range of data characteristics. Our findings show that no single method consistently dominates across all settings; instead, the effectiveness of imbalanced learning methods depends strongly on dataset characteristics and computational constraints. Based on these findings, we provide practical recommendations for selecting appropriate methods in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14915
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TILBench: A Systematic Benchmark for Tabular Imbalanced Learning Across Data Regimes
Liu, Ruizhe
Luo, Jiaqi
Machine Learning
Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed algorithms, a systematic empirical understanding of how different imbalanced learning methods behave across diverse data characteristics is still lacking. In particular, it remains unclear how different method families compare in predictive performance, robustness under varying data characteristics, and computational scalability. In this work, we present Tabular Imbalanced Learning Benchmark (TILBench), a large-scale empirical benchmark for tabular imbalanced learning. TILBench evaluates more than 40 representative algorithms across 57 diverse tabular datasets, resulting in over 200000 controlled experiments across a wide range of data characteristics. Our findings show that no single method consistently dominates across all settings; instead, the effectiveness of imbalanced learning methods depends strongly on dataset characteristics and computational constraints. Based on these findings, we provide practical recommendations for selecting appropriate methods in real-world applications.
title TILBench: A Systematic Benchmark for Tabular Imbalanced Learning Across Data Regimes
topic Machine Learning
url https://arxiv.org/abs/2605.14915