Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Firdoussi, Aymane El, Seddik, Mohamed El Amine, Hayou, Soufiane, Alami, Reda, Alzubaidi, Ahmed, Hacid, Hakim
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917800869625856
author Firdoussi, Aymane El
Seddik, Mohamed El Amine
Hayou, Soufiane
Alami, Reda
Alzubaidi, Ahmed
Hacid, Hakim
author_facet Firdoussi, Aymane El
Seddik, Mohamed El Amine
Hayou, Soufiane
Alami, Reda
Alzubaidi, Ahmed
Hacid, Hakim
contents Synthetic data has gained attention for training large language models, but poor-quality data can harm performance (see, e.g., Shumailov et al. (2023); Seddik et al. (2024)). A potential solution is data pruning, which retains only high-quality data based on a score function (human or machine feedback). Previous work Feng et al. (2024) analyzed models trained on synthetic data as sample size increases. We extend this by using random matrix theory to derive the performance of a binary classifier trained on a mix of real and pruned synthetic data in a high dimensional setting. Our findings identify conditions where synthetic data could improve performance, focusing on the quality of the generative model and verification strategy. We also show a smooth phase transition in synthetic label noise, contrasting with prior sharp behavior in infinite sample limits. Experiments with toy models and large language models validate our theoretical results.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08942
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
Firdoussi, Aymane El
Seddik, Mohamed El Amine
Hayou, Soufiane
Alami, Reda
Alzubaidi, Ahmed
Hacid, Hakim
Machine Learning
Artificial Intelligence
Statistics Theory
Synthetic data has gained attention for training large language models, but poor-quality data can harm performance (see, e.g., Shumailov et al. (2023); Seddik et al. (2024)). A potential solution is data pruning, which retains only high-quality data based on a score function (human or machine feedback). Previous work Feng et al. (2024) analyzed models trained on synthetic data as sample size increases. We extend this by using random matrix theory to derive the performance of a binary classifier trained on a mix of real and pruned synthetic data in a high dimensional setting. Our findings identify conditions where synthetic data could improve performance, focusing on the quality of the generative model and verification strategy. We also show a smooth phase transition in synthetic label noise, contrasting with prior sharp behavior in infinite sample limits. Experiments with toy models and large language models validate our theoretical results.
title Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
topic Machine Learning
Artificial Intelligence
Statistics Theory
url https://arxiv.org/abs/2410.08942