ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917948961062912 |
|---|---|
| author | Haji-Ali, Moayed Balakrishnan, Guha Ordonez, Vicente |
| author_facet | Haji-Ali, Moayed Balakrishnan, Guha Ordonez, Vicente |
| contents | Diffusion models have revolutionized image generation in recent years, yet they are still limited to a few sizes and aspect ratios. We propose ElasticDiffusion, a novel training-free decoding method that enables pretrained text-to-image diffusion models to generate images with various sizes. ElasticDiffusion attempts to decouple the generation trajectory of a pretrained model into local and global signals. The local signal controls low-level pixel information and can be estimated on local patches, while the global signal is used to maintain overall structural consistency and is estimated with a reference image. We test our method on CelebA-HQ (faces) and LAION-COCO (objects/indoor/outdoor scenes). Our experiments and qualitative results show superior image coherence quality across aspect ratios compared to MultiDiffusion and the standard decoding strategy of Stable Diffusion. Project page: https://elasticdiffusion.github.io/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_18822 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation Haji-Ali, Moayed Balakrishnan, Guha Ordonez, Vicente Computer Vision and Pattern Recognition Diffusion models have revolutionized image generation in recent years, yet they are still limited to a few sizes and aspect ratios. We propose ElasticDiffusion, a novel training-free decoding method that enables pretrained text-to-image diffusion models to generate images with various sizes. ElasticDiffusion attempts to decouple the generation trajectory of a pretrained model into local and global signals. The local signal controls low-level pixel information and can be estimated on local patches, while the global signal is used to maintain overall structural consistency and is estimated with a reference image. We test our method on CelebA-HQ (faces) and LAION-COCO (objects/indoor/outdoor scenes). Our experiments and qualitative results show superior image coherence quality across aspect ratios compared to MultiDiffusion and the standard decoding strategy of Stable Diffusion. Project page: https://elasticdiffusion.github.io/ |
| title | ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2311.18822 |