Synthetic Oversampling: Theory and A Practical Approach Using LLMs to Address Data Imbalance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nakada, Ryumei, Xu, Yichen, Li, Lexin, Zhang, Linjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910014735646720
author Nakada, Ryumei
Xu, Yichen
Li, Lexin
Zhang, Linjun
author_facet Nakada, Ryumei
Xu, Yichen
Li, Lexin
Zhang, Linjun
contents Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would compromise the accuracy, robustness and generalizability of the learned models. Recent advances have proposed leveraging the flexibility and generative capabilities of large language models (LLMs), typically built on transformer architectures, to generate synthetic samples and to augment the observed data. In the context of imbalanced data, LLMs are used to oversample underrepresented groups and have shown promising improvements. However, there is a clear lack of theoretical understanding of such synthetic data approaches. In this article, we develop novel theoretical foundations to systematically study the roles of synthetic samples in addressing imbalanced classification and spurious correlation. Specifically, we first explicitly quantify the benefits of synthetic oversampling. Next, we analyze the scaling dynamics in synthetic data augmentation, and derive the corresponding scaling law. Finally, we demonstrate the capacity of transformer models to generate high-quality synthetic samples. We further conduct extensive numerical experiments to validate the efficacy of the LLM-based synthetic oversampling and augmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03628
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Synthetic Oversampling: Theory and A Practical Approach Using LLMs to Address Data Imbalance
Nakada, Ryumei
Xu, Yichen
Li, Lexin
Zhang, Linjun
Machine Learning
Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would compromise the accuracy, robustness and generalizability of the learned models. Recent advances have proposed leveraging the flexibility and generative capabilities of large language models (LLMs), typically built on transformer architectures, to generate synthetic samples and to augment the observed data. In the context of imbalanced data, LLMs are used to oversample underrepresented groups and have shown promising improvements. However, there is a clear lack of theoretical understanding of such synthetic data approaches. In this article, we develop novel theoretical foundations to systematically study the roles of synthetic samples in addressing imbalanced classification and spurious correlation. Specifically, we first explicitly quantify the benefits of synthetic oversampling. Next, we analyze the scaling dynamics in synthetic data augmentation, and derive the corresponding scaling law. Finally, we demonstrate the capacity of transformer models to generate high-quality synthetic samples. We further conduct extensive numerical experiments to validate the efficacy of the LLM-based synthetic oversampling and augmentation.
title Synthetic Oversampling: Theory and A Practical Approach Using LLMs to Address Data Imbalance
topic Machine Learning
url https://arxiv.org/abs/2406.03628