Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Zizhao, Rostami, Mohammad, Thomason, Jesse
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910941899128832
author Hu, Zizhao
Rostami, Mohammad
Thomason, Jesse
author_facet Hu, Zizhao
Rostami, Mohammad
Thomason, Jesse
contents Recent research has highlighted the risk of generative model collapse, where performance progressively degrades when continually trained on self-generated data. However, existing exploration on model collapse is limited to single, unimodal models, limiting our understanding in more realistic scenarios, such as diverse multi-modal AI agents interacting autonomously through synthetic data and continually evolving. We expand the synthetic data training and model collapse study to multi-modal vision-language generative systems, such as vision-language models (VLMs) and text-to-image diffusion models, as well as recursive generate-train loops with multiple models. We find that model collapse, previously observed in single-modality generative models, exhibits distinct characteristics in the multi-modal context, such as improved vision-language alignment and increased variance in VLM image-captioning task. Additionally, we find that general approaches such as increased decoding budgets, greater model diversity, and relabeling with frozen models can effectively mitigate model collapse. Our findings provide initial insights and practical guidelines for reducing the risk of model collapse in self-improving multi-agent AI systems and curating robust multi-modal synthetic datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
Hu, Zizhao
Rostami, Mohammad
Thomason, Jesse
Machine Learning
Artificial Intelligence
Recent research has highlighted the risk of generative model collapse, where performance progressively degrades when continually trained on self-generated data. However, existing exploration on model collapse is limited to single, unimodal models, limiting our understanding in more realistic scenarios, such as diverse multi-modal AI agents interacting autonomously through synthetic data and continually evolving. We expand the synthetic data training and model collapse study to multi-modal vision-language generative systems, such as vision-language models (VLMs) and text-to-image diffusion models, as well as recursive generate-train loops with multiple models. We find that model collapse, previously observed in single-modality generative models, exhibits distinct characteristics in the multi-modal context, such as improved vision-language alignment and increased variance in VLM image-captioning task. Additionally, we find that general approaches such as increased decoding budgets, greater model diversity, and relabeling with frozen models can effectively mitigate model collapse. Our findings provide initial insights and practical guidelines for reducing the risk of model collapse in self-improving multi-agent AI systems and curating robust multi-modal synthetic datasets.
title Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.08803