Can Diffusion Models Disentangle? A Theoretical Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911176327168000 |
|---|---|
| author | Wang, Liming Mirza, Muhammad Jehanzeb Gong, Yishu Gong, Yuan Zhang, Jiaqi Tracey, Brian H. Placek, Katerina Vilela, Marco Glass, James R. |
| author_facet | Wang, Liming Mirza, Muhammad Jehanzeb Gong, Yishu Gong, Yuan Zhang, Jiaqi Tracey, Brian H. Placek, Katerina Vilela, Marco Glass, James R. |
| contents | This paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations. Within this framework, we establish identifiability conditions for general disentangled latent variable models, analyze training dynamics, and derive sample complexity bounds for disentangled latent subspace models. To validate our theory, we conduct disentanglement experiments across diverse tasks and modalities, including subspace recovery in latent subspace Gaussian mixture models, image colorization, image denoising, and voice conversion for speech classification. Additionally, our experiments show that training strategies inspired by our theory, such as style guidance regularization, consistently enhance disentanglement performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_00220 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Can Diffusion Models Disentangle? A Theoretical Perspective Wang, Liming Mirza, Muhammad Jehanzeb Gong, Yishu Gong, Yuan Zhang, Jiaqi Tracey, Brian H. Placek, Katerina Vilela, Marco Glass, James R. Machine Learning Artificial Intelligence Computer Vision and Pattern Recognition This paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations. Within this framework, we establish identifiability conditions for general disentangled latent variable models, analyze training dynamics, and derive sample complexity bounds for disentangled latent subspace models. To validate our theory, we conduct disentanglement experiments across diverse tasks and modalities, including subspace recovery in latent subspace Gaussian mixture models, image colorization, image denoising, and voice conversion for speech classification. Additionally, our experiments show that training strategies inspired by our theory, such as style guidance regularization, consistently enhance disentanglement performance. |
| title | Can Diffusion Models Disentangle? A Theoretical Perspective |
| topic | Machine Learning Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.00220 |