Synthetic Information towards Maximum Posterior Ratio for deep learning on Imbalanced Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Hung, Chang, Morris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917595944321024
author Nguyen, Hung
Chang, Morris
author_facet Nguyen, Hung
Chang, Morris
contents This study examines the impact of class-imbalanced data on deep learning models and proposes a technique for data balancing by generating synthetic data for the minority class. Unlike random-based oversampling, our method prioritizes balancing the informative regions by identifying high entropy samples. Generating well-placed synthetic data can enhance machine learning algorithms accuracy and efficiency, whereas poorly-placed ones may lead to higher misclassification rates. We introduce an algorithm that maximizes the probability of generating a synthetic sample in the correct region of its class by optimizing the class posterior ratio. Additionally, to maintain data topology, synthetic data are generated within each minority sample's neighborhood. Our experimental results on forty-one datasets demonstrate the superior performance of our technique in enhancing deep-learning models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02591
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Synthetic Information towards Maximum Posterior Ratio for deep learning on Imbalanced Data
Nguyen, Hung
Chang, Morris
Machine Learning
Artificial Intelligence
This study examines the impact of class-imbalanced data on deep learning models and proposes a technique for data balancing by generating synthetic data for the minority class. Unlike random-based oversampling, our method prioritizes balancing the informative regions by identifying high entropy samples. Generating well-placed synthetic data can enhance machine learning algorithms accuracy and efficiency, whereas poorly-placed ones may lead to higher misclassification rates. We introduce an algorithm that maximizes the probability of generating a synthetic sample in the correct region of its class by optimizing the class posterior ratio. Additionally, to maintain data topology, synthetic data are generated within each minority sample's neighborhood. Our experimental results on forty-one datasets demonstrate the superior performance of our technique in enhancing deep-learning models.
title Synthetic Information towards Maximum Posterior Ratio for deep learning on Imbalanced Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2401.02591