A Comprehensive Survey on Imbalanced Data Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Xinyi, Xie, Dongting, Zhang, Yihang, Wang, Zhengren, Chen, Chong, He, Conghui, Yin, Hongzhi, Zhang, Wentao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914032861052928
author Gao, Xinyi
Xie, Dongting
Zhang, Yihang
Wang, Zhengren
Chen, Chong
He, Conghui
Yin, Hongzhi
Zhang, Wentao
author_facet Gao, Xinyi
Xie, Dongting
Zhang, Yihang
Wang, Zhengren
Chen, Chong
He, Conghui
Yin, Hongzhi
Zhang, Wentao
contents With the expansion of data availability, machine learning (ML) has achieved remarkable breakthroughs in both academia and industry. However, imbalanced data distributions are prevalent in various types of raw data and severely hinder the performance of ML by biasing the decision-making processes. To deepen the understanding of imbalanced data and facilitate the related research and applications, this survey systematically analyzes various real-world data formats and concludes existing researches for different data formats into four distinct categories: data re-balancing, feature representation, training strategy, and ensemble learning. This structured analysis helps researchers comprehensively understand the pervasive nature of imbalance across diverse data formats, thereby paving a clearer path toward achieving specific research goals. We provide an overview of relevant open-source libraries, spotlight current challenges, and offer novel insights aimed at fostering future advancements in this critical area of study.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08960
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comprehensive Survey on Imbalanced Data Learning
Gao, Xinyi
Xie, Dongting
Zhang, Yihang
Wang, Zhengren
Chen, Chong
He, Conghui
Yin, Hongzhi
Zhang, Wentao
Machine Learning
With the expansion of data availability, machine learning (ML) has achieved remarkable breakthroughs in both academia and industry. However, imbalanced data distributions are prevalent in various types of raw data and severely hinder the performance of ML by biasing the decision-making processes. To deepen the understanding of imbalanced data and facilitate the related research and applications, this survey systematically analyzes various real-world data formats and concludes existing researches for different data formats into four distinct categories: data re-balancing, feature representation, training strategy, and ensemble learning. This structured analysis helps researchers comprehensively understand the pervasive nature of imbalance across diverse data formats, thereby paving a clearer path toward achieving specific research goals. We provide an overview of relevant open-source libraries, spotlight current challenges, and offer novel insights aimed at fostering future advancements in this critical area of study.
title A Comprehensive Survey on Imbalanced Data Learning
topic Machine Learning
url https://arxiv.org/abs/2502.08960