Less is More: Adaptive Coverage for Synthetic Training Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tavakkol, Sasan, Springer, Max, Bateni, Mohammadhossein, Bulut, Neslihan, Cohen-Addad, Vincent, Hajiaghayi, MohammadTaghi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908465202462720
author Tavakkol, Sasan
Springer, Max
Bateni, Mohammadhossein
Bulut, Neslihan
Cohen-Addad, Vincent
Hajiaghayi, MohammadTaghi
author_facet Tavakkol, Sasan
Springer, Max
Bateni, Mohammadhossein
Bulut, Neslihan
Cohen-Addad, Vincent
Hajiaghayi, MohammadTaghi
contents Synthetic training data generation with Large Language Models (LLMs) like Google's Gemma and OpenAI's GPT offer a promising solution to the challenge of obtaining large, labeled datasets for training classifiers. When rapid model deployment is critical, such as in classifying emerging social media trends or combating new forms of online abuse tied to current events, the ability to generate training data is invaluable. While prior research has examined the comparability of synthetic data to human-labeled data, this study introduces a novel sampling algorithm, based on the maximum coverage problem, to select a representative subset from a synthetically generated dataset. Our results demonstrate that training a classifier on this contextually sampled subset achieves superior performance compared to training on the entire dataset. This "less is more" approach not only improves model accuracy but also reduces the volume of data required, leading to potentially more efficient model fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Less is More: Adaptive Coverage for Synthetic Training Data
Tavakkol, Sasan
Springer, Max
Bateni, Mohammadhossein
Bulut, Neslihan
Cohen-Addad, Vincent
Hajiaghayi, MohammadTaghi
Machine Learning
Synthetic training data generation with Large Language Models (LLMs) like Google's Gemma and OpenAI's GPT offer a promising solution to the challenge of obtaining large, labeled datasets for training classifiers. When rapid model deployment is critical, such as in classifying emerging social media trends or combating new forms of online abuse tied to current events, the ability to generate training data is invaluable. While prior research has examined the comparability of synthetic data to human-labeled data, this study introduces a novel sampling algorithm, based on the maximum coverage problem, to select a representative subset from a synthetically generated dataset. Our results demonstrate that training a classifier on this contextually sampled subset achieves superior performance compared to training on the entire dataset. This "less is more" approach not only improves model accuracy but also reduces the volume of data required, leading to potentially more efficient model fine-tuning.
title Less is More: Adaptive Coverage for Synthetic Training Data
topic Machine Learning
url https://arxiv.org/abs/2504.14508