Self-Correcting Self-Consuming Loops for Generative Model Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gillman, Nate, Freeman, Michael, Aggarwal, Daksh, Hsu, Chia-Hong, Luo, Calvin, Tian, Yonglong, Sun, Chen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911911395721216
author Gillman, Nate
Freeman, Michael
Aggarwal, Daksh
Hsu, Chia-Hong
Luo, Calvin
Tian, Yonglong
Sun, Chen
author_facet Gillman, Nate
Freeman, Michael
Aggarwal, Daksh
Hsu, Chia-Hong
Luo, Calvin
Tian, Yonglong
Sun, Chen
contents As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic data for representation learning, using synthetic data for generative model training creates "self-consuming loops" which may lead to training instability or even collapse, unless certain conditions are met. Our paper aims to stabilize self-consuming generative model training. Our theoretical results demonstrate that by introducing an idealized correction function, which maps a data point to be more likely under the true data distribution, self-consuming loops can be made exponentially more stable. We then propose self-correction functions, which rely on expert knowledge (e.g. the laws of physics programmed in a simulator), and aim to approximate the idealized corrector automatically and at scale. We empirically validate the effectiveness of self-correcting self-consuming loops on the challenging human motion synthesis task, and observe that it successfully avoids model collapse, even when the ratio of synthetic data to real data is as high as 100%.
format Preprint
id arxiv_https___arxiv_org_abs_2402_07087
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Correcting Self-Consuming Loops for Generative Model Training
Gillman, Nate
Freeman, Michael
Aggarwal, Daksh
Hsu, Chia-Hong
Luo, Calvin
Tian, Yonglong
Sun, Chen
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic data for representation learning, using synthetic data for generative model training creates "self-consuming loops" which may lead to training instability or even collapse, unless certain conditions are met. Our paper aims to stabilize self-consuming generative model training. Our theoretical results demonstrate that by introducing an idealized correction function, which maps a data point to be more likely under the true data distribution, self-consuming loops can be made exponentially more stable. We then propose self-correction functions, which rely on expert knowledge (e.g. the laws of physics programmed in a simulator), and aim to approximate the idealized corrector automatically and at scale. We empirically validate the effectiveness of self-correcting self-consuming loops on the challenging human motion synthesis task, and observe that it successfully avoids model collapse, even when the ratio of synthetic data to real data is as high as 100%.
title Self-Correcting Self-Consuming Loops for Generative Model Training
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.07087