Diffusion Self-Distillation for Zero-Shot Customized Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912135120945152 |
|---|---|
| author | Cai, Shengqu Chan, Eric Zhang, Yunzhi Guibas, Leonidas Wu, Jiajun Wetzstein, Gordon |
| author_facet | Cai, Shengqu Chan, Eric Zhang, Yunzhi Guibas, Leonidas Wu, Jiajun Wetzstein, Gordon |
| contents | Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts, i.e., "identity-preserving generation". This setting, along with many other tasks (e.g., relighting), is a natural fit for image+text-conditional generative models. However, there is insufficient high-quality paired data to train such a model directly. We propose Diffusion Self-Distillation, a method for using a pre-trained text-to-image model to generate its own dataset for text-conditioned image-to-image tasks. We first leverage a text-to-image diffusion model's in-context generation ability to create grids of images and curate a large paired dataset with the help of a Visual-Language Model. We then fine-tune the text-to-image model into a text+image-to-image model using the curated paired dataset. We demonstrate that Diffusion Self-Distillation outperforms existing zero-shot methods and is competitive with per-instance tuning techniques on a wide range of identity-preservation generation tasks, without requiring test-time optimization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_18616 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Diffusion Self-Distillation for Zero-Shot Customized Image Generation Cai, Shengqu Chan, Eric Zhang, Yunzhi Guibas, Leonidas Wu, Jiajun Wetzstein, Gordon Computer Vision and Pattern Recognition Artificial Intelligence Graphics Machine Learning Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts, i.e., "identity-preserving generation". This setting, along with many other tasks (e.g., relighting), is a natural fit for image+text-conditional generative models. However, there is insufficient high-quality paired data to train such a model directly. We propose Diffusion Self-Distillation, a method for using a pre-trained text-to-image model to generate its own dataset for text-conditioned image-to-image tasks. We first leverage a text-to-image diffusion model's in-context generation ability to create grids of images and curate a large paired dataset with the help of a Visual-Language Model. We then fine-tune the text-to-image model into a text+image-to-image model using the curated paired dataset. We demonstrate that Diffusion Self-Distillation outperforms existing zero-shot methods and is competitive with per-instance tuning techniques on a wide range of identity-preservation generation tasks, without requiring test-time optimization. |
| title | Diffusion Self-Distillation for Zero-Shot Customized Image Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics Machine Learning |
| url | https://arxiv.org/abs/2411.18616 |