Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zizhao, Rostami, Mohammad, Thomason, Jesse |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
by: Hu, Zizhao, et al.
Published: (2026)
by: Hu, Zizhao, et al.
Published: (2026)
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
by: Hu, Zizhao, et al.
Published: (2026)
by: Hu, Zizhao, et al.
Published: (2026)
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
by: Cai, Yuliang, et al.
Published: (2025)
by: Cai, Yuliang, et al.
Published: (2025)
M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
An Intermediate Fusion ViT Enables Efficient Text-Image Alignment in Diffusion Models
by: Hu, Zizhao, et al.
Published: (2024)
by: Hu, Zizhao, et al.
Published: (2024)
Self-Improving Diffusion Models with Synthetic Data
by: Alemohammad, Sina, et al.
Published: (2024)
by: Alemohammad, Sina, et al.
Published: (2024)
CodecLM: Aligning Language Models with Tailored Synthetic Data
by: Wang, Zifeng, et al.
Published: (2024)
by: Wang, Zifeng, et al.
Published: (2024)
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
by: Falahati, Ali, et al.
Published: (2026)
by: Falahati, Ali, et al.
Published: (2026)
Synthetic Data from Diffusion Models Improve Drug Discovery Prediction
by: Hu, Bing, et al.
Published: (2024)
by: Hu, Bing, et al.
Published: (2024)
Why Do Some Inputs Break Low-Bit LLM Quantization?
by: Chang, Ting-Yun, et al.
Published: (2025)
by: Chang, Ting-Yun, et al.
Published: (2025)
Synthetic Data in Education: Empirical Insights from Traditional Resampling and Deep Generative Models
by: Chinodakufa, Tapiwa Amion, et al.
Published: (2026)
by: Chinodakufa, Tapiwa Amion, et al.
Published: (2026)
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Dominating vs. Dominated: Generative Collapse in Diffusion Models
by: Jeong, Hayeon, et al.
Published: (2025)
by: Jeong, Hayeon, et al.
Published: (2025)
Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection
by: Anwar, Abrar, et al.
Published: (2025)
by: Anwar, Abrar, et al.
Published: (2025)
Generating Synthetic Net Load Data with Physics-informed Diffusion Model
by: Zhang, Shaorong, et al.
Published: (2024)
by: Zhang, Shaorong, et al.
Published: (2024)
Hybrid Learners Do Not Forget: A Brain-Inspired Neuro-Symbolic Approach to Continual Learning
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
by: Amin, Kareem, et al.
Published: (2025)
by: Amin, Kareem, et al.
Published: (2025)
Downstream Task-Oriented Generative Model Selections on Synthetic Data Training for Fraud Detection Models
by: Cheng, Yinan, et al.
Published: (2024)
by: Cheng, Yinan, et al.
Published: (2024)
Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models
by: Lin, Jinxu, et al.
Published: (2024)
by: Lin, Jinxu, et al.
Published: (2024)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
by: Kazdan, Joshua, et al.
Published: (2024)
by: Kazdan, Joshua, et al.
Published: (2024)
SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis
by: Brockschmidt, Marie, et al.
Published: (2025)
by: Brockschmidt, Marie, et al.
Published: (2025)
A Note on Shumailov et al. (2024): `AI Models Collapse When Trained on Recursively Generated Data'
by: Borji, Ali
Published: (2024)
by: Borji, Ali
Published: (2024)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
by: Gerstgrasser, Matthias, et al.
Published: (2024)
by: Gerstgrasser, Matthias, et al.
Published: (2024)
Does Training on Synthetic Data Make Models Less Robust?
by: Zhang, Lingze, et al.
Published: (2025)
by: Zhang, Lingze, et al.
Published: (2025)
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
by: Mehta, Sushant
Published: (2025)
by: Mehta, Sushant
Published: (2025)
Synthetic Power Flow Data Generation Using Physics-Informed Denoising Diffusion Probabilistic Models
by: Wang, Junfei, et al.
Published: (2025)
by: Wang, Junfei, et al.
Published: (2025)
Creating Artificial Students that Never Existed: Leveraging Large Language Models and CTGANs for Synthetic Data Generation
by: Khalil, Mohammad, et al.
Published: (2025)
by: Khalil, Mohammad, et al.
Published: (2025)
Enhancing Pre-Trained Model-Based Class-Incremental Learning through Neural Collapse
by: He, Kun, et al.
Published: (2025)
by: He, Kun, et al.
Published: (2025)
A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
by: Fu, Shi, et al.
Published: (2025)
by: Fu, Shi, et al.
Published: (2025)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
by: Feng, Yunzhen, et al.
Published: (2024)
by: Feng, Yunzhen, et al.
Published: (2024)
A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion Models
by: Kamkari, Hamidreza, et al.
Published: (2024)
by: Kamkari, Hamidreza, et al.
Published: (2024)
TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
Ambient Diffusion Omni: Training Good Models with Bad Data
by: Daras, Giannis, et al.
Published: (2025)
by: Daras, Giannis, et al.
Published: (2025)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
by: Prakash, Hari K., et al.
Published: (2025)
by: Prakash, Hari K., et al.
Published: (2025)
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
by: Murata, Naoki, et al.
Published: (2026)
by: Murata, Naoki, et al.
Published: (2026)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024)
by: Pumacay, Wilbert, et al.
Published: (2024)
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Similar Items
-
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
by: Hu, Zizhao, et al.
Published: (2026) -
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
by: Hu, Zizhao, et al.
Published: (2026) -
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
by: Cai, Yuliang, et al.
Published: (2025) -
M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention
by: Tang, Yiming, et al.
Published: (2025) -
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)