Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Wei, Han, Andi, Bai, Mingyuan, Zhou, Huanjian, Zhang, Qixin, Suzuki, Taiji, Fukumizu, Kenji
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911698807422976
author Huang, Wei
Han, Andi
Bai, Mingyuan
Zhou, Huanjian
Zhang, Qixin
Suzuki, Taiji
Fukumizu, Kenji
author_facet Huang, Wei
Han, Andi
Bai, Mingyuan
Zhou, Huanjian
Zhang, Qixin
Suzuki, Taiji
Fukumizu, Kenji
contents Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when data is supported on low-dimensional manifolds, remains theoretically unexplained. We identify a collapse-and-refine mechanism driven by the geometry of the score function itself: at small noise scales, the diverging singularity of the score drives a rapid dimensional collapse of the induced denoising map onto the data manifold projection; at moderate noise scales, training refines the intrinsic density on the learned manifold. We instantiate this principle as Score-induced Latent Diffusion (SiLD), a two-stage framework in which both manifold learning and density estimation emerge from a single denoising score matching objective, replacing the heuristic KL regularization of VAE-based latent diffusion models. We prove that the resulting sample complexity depends on the intrinsic dimension rather than the ambient dimension. Experiments on Stacked MNIST, CelebA variants, and molecular generation benchmarks show that SiLD matches or outperforms VAE-based LDMs in generation quality and consistently improves reconstruction, validating our theoretical predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20235
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
Huang, Wei
Han, Andi
Bai, Mingyuan
Zhou, Huanjian
Zhang, Qixin
Suzuki, Taiji
Fukumizu, Kenji
Machine Learning
Artificial Intelligence
Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when data is supported on low-dimensional manifolds, remains theoretically unexplained. We identify a collapse-and-refine mechanism driven by the geometry of the score function itself: at small noise scales, the diverging singularity of the score drives a rapid dimensional collapse of the induced denoising map onto the data manifold projection; at moderate noise scales, training refines the intrinsic density on the learned manifold. We instantiate this principle as Score-induced Latent Diffusion (SiLD), a two-stage framework in which both manifold learning and density estimation emerge from a single denoising score matching objective, replacing the heuristic KL regularization of VAE-based latent diffusion models. We prove that the resulting sample complexity depends on the intrinsic dimension rather than the ambient dimension. Experiments on Stacked MNIST, CelebA variants, and molecular generation benchmarks show that SiLD matches or outperforms VAE-based LDMs in generation quality and consistently improves reconstruction, validating our theoretical predictions.
title Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.20235