Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911698807422976 |
|---|---|
| author | Huang, Wei Han, Andi Bai, Mingyuan Zhou, Huanjian Zhang, Qixin Suzuki, Taiji Fukumizu, Kenji |
| author_facet | Huang, Wei Han, Andi Bai, Mingyuan Zhou, Huanjian Zhang, Qixin Suzuki, Taiji Fukumizu, Kenji |
| contents | Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when data is supported on low-dimensional manifolds, remains theoretically unexplained. We identify a collapse-and-refine mechanism driven by the geometry of the score function itself: at small noise scales, the diverging singularity of the score drives a rapid dimensional collapse of the induced denoising map onto the data manifold projection; at moderate noise scales, training refines the intrinsic density on the learned manifold. We instantiate this principle as Score-induced Latent Diffusion (SiLD), a two-stage framework in which both manifold learning and density estimation emerge from a single denoising score matching objective, replacing the heuristic KL regularization of VAE-based latent diffusion models. We prove that the resulting sample complexity depends on the intrinsic dimension rather than the ambient dimension. Experiments on Stacked MNIST, CelebA variants, and molecular generation benchmarks show that SiLD matches or outperforms VAE-based LDMs in generation quality and consistently improves reconstruction, validating our theoretical predictions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_20235 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine Huang, Wei Han, Andi Bai, Mingyuan Zhou, Huanjian Zhang, Qixin Suzuki, Taiji Fukumizu, Kenji Machine Learning Artificial Intelligence Diffusion models generate high-dimensional data with remarkable quality, yet how their training efficiently learns the score function, bypassing the curse of dimensionality when data is supported on low-dimensional manifolds, remains theoretically unexplained. We identify a collapse-and-refine mechanism driven by the geometry of the score function itself: at small noise scales, the diverging singularity of the score drives a rapid dimensional collapse of the induced denoising map onto the data manifold projection; at moderate noise scales, training refines the intrinsic density on the learned manifold. We instantiate this principle as Score-induced Latent Diffusion (SiLD), a two-stage framework in which both manifold learning and density estimation emerge from a single denoising score matching objective, replacing the heuristic KL regularization of VAE-based latent diffusion models. We prove that the resulting sample complexity depends on the intrinsic dimension rather than the ambient dimension. Experiments on Stacked MNIST, CelebA variants, and molecular generation benchmarks show that SiLD matches or outperforms VAE-based LDMs in generation quality and consistently improves reconstruction, validating our theoretical predictions. |
| title | Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2605.20235 |