Learning Confidence Bounds for Classification with Imbalanced Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Clifford, Matt, Erskine, Jonathan, Hepburn, Alexander, Santos-Rodríguez, Raúl, Garcia-Garcia, Dario
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912662836740096
author Clifford, Matt
Erskine, Jonathan
Hepburn, Alexander
Santos-Rodríguez, Raúl
Garcia-Garcia, Dario
author_facet Clifford, Matt
Erskine, Jonathan
Hepburn, Alexander
Santos-Rodríguez, Raúl
Garcia-Garcia, Dario
contents Class imbalance poses a significant challenge in classification tasks, where traditional approaches often lead to biased models and unreliable predictions. Undersampling and oversampling techniques have been commonly employed to address this issue, yet they suffer from inherent limitations stemming from their simplistic approach such as loss of information and additional biases respectively. In this paper, we propose a novel framework that leverages learning theory and concentration inequalities to overcome the shortcomings of traditional solutions. We focus on understanding the uncertainty in a class-dependent manner, as captured by confidence bounds that we directly embed into the learning process. By incorporating class-dependent estimates, our method can effectively adapt to the varying degrees of imbalance across different classes, resulting in more robust and reliable classification outcomes. We empirically show how our framework provides a promising direction for handling imbalanced data in classification tasks, offering practitioners a valuable tool for building more accurate and trustworthy models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11878
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Confidence Bounds for Classification with Imbalanced Data
Clifford, Matt
Erskine, Jonathan
Hepburn, Alexander
Santos-Rodríguez, Raúl
Garcia-Garcia, Dario
Machine Learning
Class imbalance poses a significant challenge in classification tasks, where traditional approaches often lead to biased models and unreliable predictions. Undersampling and oversampling techniques have been commonly employed to address this issue, yet they suffer from inherent limitations stemming from their simplistic approach such as loss of information and additional biases respectively. In this paper, we propose a novel framework that leverages learning theory and concentration inequalities to overcome the shortcomings of traditional solutions. We focus on understanding the uncertainty in a class-dependent manner, as captured by confidence bounds that we directly embed into the learning process. By incorporating class-dependent estimates, our method can effectively adapt to the varying degrees of imbalance across different classes, resulting in more robust and reliable classification outcomes. We empirically show how our framework provides a promising direction for handling imbalanced data in classification tasks, offering practitioners a valuable tool for building more accurate and trustworthy models.
title Learning Confidence Bounds for Classification with Imbalanced Data
topic Machine Learning
url https://arxiv.org/abs/2407.11878