Oversampling and Downsampling with Core-Boundary Awareness: A Data Quality-Driven Approach

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Belhaouari, Samir Brahim, Kahalan, Yunis Carreon, Shaffique, Humaira, Belhaouari, Ismael, Islam, Ashhadul
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908556609978368
author Belhaouari, Samir Brahim
Kahalan, Yunis Carreon
Shaffique, Humaira
Belhaouari, Ismael
Islam, Ashhadul
author_facet Belhaouari, Samir Brahim
Kahalan, Yunis Carreon
Shaffique, Humaira
Belhaouari, Ismael
Islam, Ashhadul
contents The effectiveness of machine learning models, particularly in unbalanced classification tasks, is often hindered by the failure to differentiate between critical instances near the decision boundary and redundant samples concentrated in the core of the data distribution. In this paper, we propose a method to systematically identify and differentiate between these two types of data. Through extensive experiments on multiple benchmark datasets, we show that the boundary data oversampling method improves the F1 score by up to 10\% on 96\% of the datasets, whereas our core-aware reduction method compresses datasets up to 90\% while preserving their accuracy, making it 10 times more powerful than the original dataset. Beyond imbalanced classification, our method has broader implications for efficient model training, particularly in computationally expensive domains such as Large Language Model (LLM) training. By prioritizing high-quality, decision-relevant data, our approach can be extended to text, multimodal, and self-supervised learning scenarios, offering a pathway to faster convergence, improved generalization, and significant computational savings. This work paves the way for future research in data-efficient learning, where intelligent sampling replaces brute-force expansion, driving the next generation of AI advancements. Our code is available as a Python package at https://pypi.org/project/adaptive-resampling/ .
format Preprint
id arxiv_https___arxiv_org_abs_2509_19856
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Oversampling and Downsampling with Core-Boundary Awareness: A Data Quality-Driven Approach
Belhaouari, Samir Brahim
Kahalan, Yunis Carreon
Shaffique, Humaira
Belhaouari, Ismael
Islam, Ashhadul
Machine Learning
The effectiveness of machine learning models, particularly in unbalanced classification tasks, is often hindered by the failure to differentiate between critical instances near the decision boundary and redundant samples concentrated in the core of the data distribution. In this paper, we propose a method to systematically identify and differentiate between these two types of data. Through extensive experiments on multiple benchmark datasets, we show that the boundary data oversampling method improves the F1 score by up to 10\% on 96\% of the datasets, whereas our core-aware reduction method compresses datasets up to 90\% while preserving their accuracy, making it 10 times more powerful than the original dataset. Beyond imbalanced classification, our method has broader implications for efficient model training, particularly in computationally expensive domains such as Large Language Model (LLM) training. By prioritizing high-quality, decision-relevant data, our approach can be extended to text, multimodal, and self-supervised learning scenarios, offering a pathway to faster convergence, improved generalization, and significant computational savings. This work paves the way for future research in data-efficient learning, where intelligent sampling replaces brute-force expansion, driving the next generation of AI advancements. Our code is available as a Python package at https://pypi.org/project/adaptive-resampling/ .
title Oversampling and Downsampling with Core-Boundary Awareness: A Data Quality-Driven Approach
topic Machine Learning
url https://arxiv.org/abs/2509.19856