Saved in:
Bibliographic Details
Main Authors: Ashraf, Imran, Ullah, Mukhtar, Nadeem, Muhammad Faisal, Noor, Muhammad Nouman
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.21380
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916970666917888
author Ashraf, Imran
Ullah, Mukhtar
Nadeem, Muhammad Faisal
Noor, Muhammad Nouman
author_facet Ashraf, Imran
Ullah, Mukhtar
Nadeem, Muhammad Faisal
Noor, Muhammad Nouman
contents Deep Learning models have transformed various domains, including the healthcare sector, particularly biomedical image classification by learning intricate features and enabling accurate diagnostics pertaining to complex diseases. Recent studies have adopted two different approaches to train DL models: training from scratch and transfer learning. Both approaches demand substantial computational time and resources due to the involvement of massive datasets in model training. These computational demands are further increased due to the design-space exploration required for selecting optimal hyperparameters, which typically necessitates several training rounds. With the growing sizes of datasets, exploring solutions to this problem has recently gained the research community's attention. A plausible solution is to select a subset of the dataset for training and hyperparameter search. This subset, referred to as the corset, must be a representative set of the original dataset. A straightforward approach to selecting the coreset could be employing random sampling, albeit at the cost of compromising the representativeness of the original dataset. A critical limitation of random sampling is the bias towards the dominant classes in an imbalanced dataset. Even if the dataset has inter-class balance, this random sampling will not capture intra-class diversity. This study addresses this issue by introducing an intelligent, lightweight mechanism for coreset selection. Specifically, it proposes a method to extract intra-class diversity, forming per-class clusters that are utilized for the final sampling. We demonstrate the efficacy of the proposed methodology by conducting extensive classification experiments on a well-known biomedical imaging dataset. Results demonstrate that the proposed scheme outperforms the random sampling approach on several performance metrics for uniform conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21380
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Coreset selection based on Intra-class diversity
Ashraf, Imran
Ullah, Mukhtar
Nadeem, Muhammad Faisal
Noor, Muhammad Nouman
Computer Vision and Pattern Recognition
Machine Learning
Deep Learning models have transformed various domains, including the healthcare sector, particularly biomedical image classification by learning intricate features and enabling accurate diagnostics pertaining to complex diseases. Recent studies have adopted two different approaches to train DL models: training from scratch and transfer learning. Both approaches demand substantial computational time and resources due to the involvement of massive datasets in model training. These computational demands are further increased due to the design-space exploration required for selecting optimal hyperparameters, which typically necessitates several training rounds. With the growing sizes of datasets, exploring solutions to this problem has recently gained the research community's attention. A plausible solution is to select a subset of the dataset for training and hyperparameter search. This subset, referred to as the corset, must be a representative set of the original dataset. A straightforward approach to selecting the coreset could be employing random sampling, albeit at the cost of compromising the representativeness of the original dataset. A critical limitation of random sampling is the bias towards the dominant classes in an imbalanced dataset. Even if the dataset has inter-class balance, this random sampling will not capture intra-class diversity. This study addresses this issue by introducing an intelligent, lightweight mechanism for coreset selection. Specifically, it proposes a method to extract intra-class diversity, forming per-class clusters that are utilized for the final sampling. We demonstrate the efficacy of the proposed methodology by conducting extensive classification experiments on a well-known biomedical imaging dataset. Results demonstrate that the proposed scheme outperforms the random sampling approach on several performance metrics for uniform conditions.
title Coreset selection based on Intra-class diversity
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.21380