Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914834178637824 |
|---|---|
| author | Vasconcelos, Cristina N. Rashwan, Abdullah Waters, Austin Walker, Trevor Xu, Keyang Yan, Jimmy Qian, Rui Luo, Shixin Parekh, Zarana Bunner, Andrew Fei, Hongliang Garg, Roopal Guo, Mandy Kajic, Ivana Li, Yeqing Nandwani, Henna Pont-Tuset, Jordi Onoe, Yasumasa Rosston, Sarah Wang, Su Zhou, Wenlei Swersky, Kevin Fleet, David J. Baldridge, Jason M. Wang, Oliver |
| author_facet | Vasconcelos, Cristina N. Rashwan, Abdullah Waters, Austin Walker, Trevor Xu, Keyang Yan, Jimmy Qian, Rui Luo, Shixin Parekh, Zarana Bunner, Andrew Fei, Hongliang Garg, Roopal Guo, Mandy Kajic, Ivana Li, Yeqing Nandwani, Henna Pont-Tuset, Jordi Onoe, Yasumasa Rosston, Sarah Wang, Su Zhou, Wenlei Swersky, Kevin Fleet, David J. Baldridge, Jason M. Wang, Oliver |
| contents | We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale, high-resolution models. without the needs for cascaded super-resolution components. The key insight stems from careful pre-training of core components, namely, those responsible for text-to-image alignment {\it vs.} high-resolution rendering. We first demonstrate the benefits of scaling a {\it Shallow UNet}, with no down(up)-sampling enc(dec)oder. Scaling its deep core layers is shown to improve alignment, object structure, and composition. Building on this core model, we propose a greedy algorithm that grows the architecture into high-resolution end-to-end models, while preserving the integrity of the pre-trained representation, stabilizing training, and reducing the need for large high-resolution datasets. This enables a single stage model capable of generating high-resolution images without the need of a super-resolution cascade. Our key results rely on public datasets and show that we are able to train non-cascaded models up to 8B parameters with no further regularization schemes. Vermeer, our full pipeline model trained with internal datasets to produce 1024x1024 images, without cascades, is preferred by 44.0% vs. 21.4% human evaluators over SDXL. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_16759 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models Vasconcelos, Cristina N. Rashwan, Abdullah Waters, Austin Walker, Trevor Xu, Keyang Yan, Jimmy Qian, Rui Luo, Shixin Parekh, Zarana Bunner, Andrew Fei, Hongliang Garg, Roopal Guo, Mandy Kajic, Ivana Li, Yeqing Nandwani, Henna Pont-Tuset, Jordi Onoe, Yasumasa Rosston, Sarah Wang, Su Zhou, Wenlei Swersky, Kevin Fleet, David J. Baldridge, Jason M. Wang, Oliver Computer Vision and Pattern Recognition Machine Learning We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale, high-resolution models. without the needs for cascaded super-resolution components. The key insight stems from careful pre-training of core components, namely, those responsible for text-to-image alignment {\it vs.} high-resolution rendering. We first demonstrate the benefits of scaling a {\it Shallow UNet}, with no down(up)-sampling enc(dec)oder. Scaling its deep core layers is shown to improve alignment, object structure, and composition. Building on this core model, we propose a greedy algorithm that grows the architecture into high-resolution end-to-end models, while preserving the integrity of the pre-trained representation, stabilizing training, and reducing the need for large high-resolution datasets. This enables a single stage model capable of generating high-resolution images without the need of a super-resolution cascade. Our key results rely on public datasets and show that we are able to train non-cascaded models up to 8B parameters with no further regularization schemes. Vermeer, our full pipeline model trained with internal datasets to produce 1024x1024 images, without cascades, is preferred by 44.0% vs. 21.4% human evaluators over SDXL. |
| title | Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2405.16759 |