Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913998250704896 |
|---|---|
| author | Gupta, Samarth Gadde, Raghudeep Chen, Rui Martinez, Aleix M. |
| author_facet | Gupta, Samarth Gadde, Raghudeep Chen, Rui Martinez, Aleix M. |
| contents | We challenge a fundamental assumption of diffusion models, namely, that a large number of latent-states or time-steps is required for training so that the reverse generative process is close to a Gaussian. We first show that with careful selection of a noise schedule, diffusion models trained over a small number of latent states (i.e. $T \sim 32$) match the performance of models trained over a much large number of latent states ($T \sim 1,000$). Second, we push this limit (on the minimum number of latent states required) to a single latent-state, which we refer to as complete disentanglement in T-space. We show that high quality samples can be easily generated by the disentangled model obtained by combining several independently trained single latent-state models. We provide extensive experiments to show that the proposed disentangled model provides 4-6$\times$ faster convergence measured across a variety of metrics on two different datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_14413 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states Gupta, Samarth Gadde, Raghudeep Chen, Rui Martinez, Aleix M. Machine Learning Computer Vision and Pattern Recognition We challenge a fundamental assumption of diffusion models, namely, that a large number of latent-states or time-steps is required for training so that the reverse generative process is close to a Gaussian. We first show that with careful selection of a noise schedule, diffusion models trained over a small number of latent states (i.e. $T \sim 32$) match the performance of models trained over a much large number of latent states ($T \sim 1,000$). Second, we push this limit (on the minimum number of latent states required) to a single latent-state, which we refer to as complete disentanglement in T-space. We show that high quality samples can be easily generated by the disentangled model obtained by combining several independently trained single latent-state models. We provide extensive experiments to show that the proposed disentangled model provides 4-6$\times$ faster convergence measured across a variety of metrics on two different datasets. |
| title | Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states |
| topic | Machine Learning Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.14413 |