Targeted synthetic data generation for tabular data via hardness characterization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912223876612096 |
|---|---|
| author | Ferracci, Tommaso Goldmann, Leonie Tabea Hinel, Anton Passino, Francesco Sanna |
| author_facet | Ferracci, Tommaso Goldmann, Leonie Tabea Hinel, Anton Passino, Francesco Sanna |
| contents | Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify beneficial and detrimental observations, we introduce a simple augmentation pipeline that generates only high-value training points based on hardness characterization, in a computationally efficient manner. We first empirically demonstrate via benchmarks on real data that Shapley-based data valuation methods perform comparably with learning-based methods in hardness characterization tasks, while offering significant computational advantages. Then, we show that synthetic data generators trained on the hardest points outperform non-targeted data augmentation on a number of tabular datasets. Our approach improves the quality of out-of-sample predictions and it is computationally more efficient compared to non-targeted methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_00759 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Targeted synthetic data generation for tabular data via hardness characterization Ferracci, Tommaso Goldmann, Leonie Tabea Hinel, Anton Passino, Francesco Sanna Machine Learning Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify beneficial and detrimental observations, we introduce a simple augmentation pipeline that generates only high-value training points based on hardness characterization, in a computationally efficient manner. We first empirically demonstrate via benchmarks on real data that Shapley-based data valuation methods perform comparably with learning-based methods in hardness characterization tasks, while offering significant computational advantages. Then, we show that synthetic data generators trained on the hardest points outperform non-targeted data augmentation on a number of tabular datasets. Our approach improves the quality of out-of-sample predictions and it is computationally more efficient compared to non-targeted methods. |
| title | Targeted synthetic data generation for tabular data via hardness characterization |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2410.00759 |