A Survey on Small Sample Imbalance Problem: Metrics, Feature Analysis, and Solutions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Shuxian, Gui, Jie, Dong, Minjing, Yu, Baosheng, Gui, Zhipeng, Dong, Lu, Tang, Yuan Yan, Kwok, James Tin-Yau
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913801151971328
author Zhao, Shuxian
Gui, Jie
Dong, Minjing
Yu, Baosheng
Gui, Zhipeng
Dong, Lu
Tang, Yuan Yan
Kwok, James Tin-Yau
author_facet Zhao, Shuxian
Gui, Jie
Dong, Minjing
Yu, Baosheng
Gui, Zhipeng
Dong, Lu
Tang, Yuan Yan
Kwok, James Tin-Yau
contents The small sample imbalance (S&I) problem is a major challenge in machine learning and data analysis. It is characterized by a small number of samples and an imbalanced class distribution, which leads to poor model performance. In addition, indistinct inter-class feature distributions further complicate classification tasks. Existing methods often rely on algorithmic heuristics without sufficiently analyzing the underlying data characteristics. We argue that a detailed analysis from the data perspective is essential before developing an appropriate solution. Therefore, this paper proposes a systematic analytical framework for the S\&I problem. We first summarize imbalance metrics and complexity analysis methods, highlighting the need for interpretable benchmarks to characterize S&I problems. Second, we review recent solutions for conventional, complexity-based, and extreme S&I problems, revealing methodological differences in handling various data distributions. Our summary finds that resampling remains a widely adopted solution. However, we conduct experiments on binary and multiclass datasets, revealing that classifier performance differences significantly exceed the improvements achieved through resampling. Finally, this paper highlights open questions and discusses future trends.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14800
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Small Sample Imbalance Problem: Metrics, Feature Analysis, and Solutions
Zhao, Shuxian
Gui, Jie
Dong, Minjing
Yu, Baosheng
Gui, Zhipeng
Dong, Lu
Tang, Yuan Yan
Kwok, James Tin-Yau
Machine Learning
Computer Vision and Pattern Recognition
The small sample imbalance (S&I) problem is a major challenge in machine learning and data analysis. It is characterized by a small number of samples and an imbalanced class distribution, which leads to poor model performance. In addition, indistinct inter-class feature distributions further complicate classification tasks. Existing methods often rely on algorithmic heuristics without sufficiently analyzing the underlying data characteristics. We argue that a detailed analysis from the data perspective is essential before developing an appropriate solution. Therefore, this paper proposes a systematic analytical framework for the S\&I problem. We first summarize imbalance metrics and complexity analysis methods, highlighting the need for interpretable benchmarks to characterize S&I problems. Second, we review recent solutions for conventional, complexity-based, and extreme S&I problems, revealing methodological differences in handling various data distributions. Our summary finds that resampling remains a widely adopted solution. However, we conduct experiments on binary and multiclass datasets, revealing that classifier performance differences significantly exceed the improvements achieved through resampling. Finally, this paper highlights open questions and discusses future trends.
title A Survey on Small Sample Imbalance Problem: Metrics, Feature Analysis, and Solutions
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.14800