Targeted synthetic data generation for tabular data via hardness characterization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferracci, Tommaso, Goldmann, Leonie Tabea, Hinel, Anton, Passino, Francesco Sanna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912223876612096
author Ferracci, Tommaso
Goldmann, Leonie Tabea
Hinel, Anton
Passino, Francesco Sanna
author_facet Ferracci, Tommaso
Goldmann, Leonie Tabea
Hinel, Anton
Passino, Francesco Sanna
contents Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify beneficial and detrimental observations, we introduce a simple augmentation pipeline that generates only high-value training points based on hardness characterization, in a computationally efficient manner. We first empirically demonstrate via benchmarks on real data that Shapley-based data valuation methods perform comparably with learning-based methods in hardness characterization tasks, while offering significant computational advantages. Then, we show that synthetic data generators trained on the hardest points outperform non-targeted data augmentation on a number of tabular datasets. Our approach improves the quality of out-of-sample predictions and it is computationally more efficient compared to non-targeted methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00759
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Targeted synthetic data generation for tabular data via hardness characterization
Ferracci, Tommaso
Goldmann, Leonie Tabea
Hinel, Anton
Passino, Francesco Sanna
Machine Learning
Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify beneficial and detrimental observations, we introduce a simple augmentation pipeline that generates only high-value training points based on hardness characterization, in a computationally efficient manner. We first empirically demonstrate via benchmarks on real data that Shapley-based data valuation methods perform comparably with learning-based methods in hardness characterization tasks, while offering significant computational advantages. Then, we show that synthetic data generators trained on the hardest points outperform non-targeted data augmentation on a number of tabular datasets. Our approach improves the quality of out-of-sample predictions and it is computationally more efficient compared to non-targeted methods.
title Targeted synthetic data generation for tabular data via hardness characterization
topic Machine Learning
url https://arxiv.org/abs/2410.00759