Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vasconcelos, Cristina N., Rashwan, Abdullah, Waters, Austin, Walker, Trevor, Xu, Keyang, Yan, Jimmy, Qian, Rui, Luo, Shixin, Parekh, Zarana, Bunner, Andrew, Fei, Hongliang, Garg, Roopal, Guo, Mandy, Kajic, Ivana, Li, Yeqing, Nandwani, Henna, Pont-Tuset, Jordi, Onoe, Yasumasa, Rosston, Sarah, Wang, Su, Zhou, Wenlei, Swersky, Kevin, Fleet, David J., Baldridge, Jason M., Wang, Oliver
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914834178637824
author Vasconcelos, Cristina N.
Rashwan, Abdullah
Waters, Austin
Walker, Trevor
Xu, Keyang
Yan, Jimmy
Qian, Rui
Luo, Shixin
Parekh, Zarana
Bunner, Andrew
Fei, Hongliang
Garg, Roopal
Guo, Mandy
Kajic, Ivana
Li, Yeqing
Nandwani, Henna
Pont-Tuset, Jordi
Onoe, Yasumasa
Rosston, Sarah
Wang, Su
Zhou, Wenlei
Swersky, Kevin
Fleet, David J.
Baldridge, Jason M.
Wang, Oliver
author_facet Vasconcelos, Cristina N.
Rashwan, Abdullah
Waters, Austin
Walker, Trevor
Xu, Keyang
Yan, Jimmy
Qian, Rui
Luo, Shixin
Parekh, Zarana
Bunner, Andrew
Fei, Hongliang
Garg, Roopal
Guo, Mandy
Kajic, Ivana
Li, Yeqing
Nandwani, Henna
Pont-Tuset, Jordi
Onoe, Yasumasa
Rosston, Sarah
Wang, Su
Zhou, Wenlei
Swersky, Kevin
Fleet, David J.
Baldridge, Jason M.
Wang, Oliver
contents We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale, high-resolution models. without the needs for cascaded super-resolution components. The key insight stems from careful pre-training of core components, namely, those responsible for text-to-image alignment {\it vs.} high-resolution rendering. We first demonstrate the benefits of scaling a {\it Shallow UNet}, with no down(up)-sampling enc(dec)oder. Scaling its deep core layers is shown to improve alignment, object structure, and composition. Building on this core model, we propose a greedy algorithm that grows the architecture into high-resolution end-to-end models, while preserving the integrity of the pre-trained representation, stabilizing training, and reducing the need for large high-resolution datasets. This enables a single stage model capable of generating high-resolution images without the need of a super-resolution cascade. Our key results rely on public datasets and show that we are able to train non-cascaded models up to 8B parameters with no further regularization schemes. Vermeer, our full pipeline model trained with internal datasets to produce 1024x1024 images, without cascades, is preferred by 44.0% vs. 21.4% human evaluators over SDXL.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16759
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
Vasconcelos, Cristina N.
Rashwan, Abdullah
Waters, Austin
Walker, Trevor
Xu, Keyang
Yan, Jimmy
Qian, Rui
Luo, Shixin
Parekh, Zarana
Bunner, Andrew
Fei, Hongliang
Garg, Roopal
Guo, Mandy
Kajic, Ivana
Li, Yeqing
Nandwani, Henna
Pont-Tuset, Jordi
Onoe, Yasumasa
Rosston, Sarah
Wang, Su
Zhou, Wenlei
Swersky, Kevin
Fleet, David J.
Baldridge, Jason M.
Wang, Oliver
Computer Vision and Pattern Recognition
Machine Learning
We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale, high-resolution models. without the needs for cascaded super-resolution components. The key insight stems from careful pre-training of core components, namely, those responsible for text-to-image alignment {\it vs.} high-resolution rendering. We first demonstrate the benefits of scaling a {\it Shallow UNet}, with no down(up)-sampling enc(dec)oder. Scaling its deep core layers is shown to improve alignment, object structure, and composition. Building on this core model, we propose a greedy algorithm that grows the architecture into high-resolution end-to-end models, while preserving the integrity of the pre-trained representation, stabilizing training, and reducing the need for large high-resolution datasets. This enables a single stage model capable of generating high-resolution images without the need of a super-resolution cascade. Our key results rely on public datasets and show that we are able to train non-cascaded models up to 8B parameters with no further regularization schemes. Vermeer, our full pipeline model trained with internal datasets to produce 1024x1024 images, without cascades, is preferred by 44.0% vs. 21.4% human evaluators over SDXL.
title Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.16759