On the Stability of Iterative Retraining of Generative Models on their own Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bertrand, Quentin, Bose, Avishek Joey, Duplessis, Alexandre, Jiralerspong, Marco, Gidel, Gauthier
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914737844912128
author Bertrand, Quentin
Bose, Avishek Joey
Duplessis, Alexandre
Jiralerspong, Marco
Gidel, Gauthier
author_facet Bertrand, Quentin
Bose, Avishek Joey
Duplessis, Alexandre
Jiralerspong, Marco
Gidel, Gauthier
contents Deep generative models have made tremendous progress in modeling complex data, often exhibiting generation quality that surpasses a typical human's ability to discern the authenticity of samples. Undeniably, a key driver of this success is enabled by the massive amounts of web-scale data consumed by these models. Due to these models' striking performance and ease of availability, the web will inevitably be increasingly populated with synthetic content. Such a fact directly implies that future iterations of generative models will be trained on both clean and artificially generated data from past models. In this paper, we develop a framework to rigorously study the impact of training generative models on mixed datasets -- from classical training on real data to self-consuming generative models trained on purely synthetic data. We first prove the stability of iterative training under the condition that the initial generative models approximate the data distribution well enough and the proportion of clean training data (w.r.t. synthetic data) is large enough. We empirically validate our theory on both synthetic and natural images by iteratively training normalizing flows and state-of-the-art diffusion models on CIFAR10 and FFHQ.
format Preprint
id arxiv_https___arxiv_org_abs_2310_00429
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On the Stability of Iterative Retraining of Generative Models on their own Data
Bertrand, Quentin
Bose, Avishek Joey
Duplessis, Alexandre
Jiralerspong, Marco
Gidel, Gauthier
Machine Learning
Deep generative models have made tremendous progress in modeling complex data, often exhibiting generation quality that surpasses a typical human's ability to discern the authenticity of samples. Undeniably, a key driver of this success is enabled by the massive amounts of web-scale data consumed by these models. Due to these models' striking performance and ease of availability, the web will inevitably be increasingly populated with synthetic content. Such a fact directly implies that future iterations of generative models will be trained on both clean and artificially generated data from past models. In this paper, we develop a framework to rigorously study the impact of training generative models on mixed datasets -- from classical training on real data to self-consuming generative models trained on purely synthetic data. We first prove the stability of iterative training under the condition that the initial generative models approximate the data distribution well enough and the proportion of clean training data (w.r.t. synthetic data) is large enough. We empirically validate our theory on both synthetic and natural images by iteratively training normalizing flows and state-of-the-art diffusion models on CIFAR10 and FFHQ.
title On the Stability of Iterative Retraining of Generative Models on their own Data
topic Machine Learning
url https://arxiv.org/abs/2310.00429