Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Crowson, Katherine, Baumann, Stefan Andreas, Birch, Alex, Abraham, Tanishq Mathew, Kaplan, Daniel Z., Shippole, Enrico
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915891195674624
author Crowson, Katherine
Baumann, Stefan Andreas
Birch, Alex
Abraham, Tanishq Mathew
Kaplan, Daniel Z.
Shippole, Enrico
author_facet Crowson, Katherine
Baumann, Stefan Andreas
Birch, Alex
Abraham, Tanishq Mathew
Kaplan, Daniel Z.
Shippole, Enrico
contents We present the Hourglass Diffusion Transformer (HDiT), an image generative model that exhibits linear scaling with pixel count, supporting training at high-resolution (e.g. $1024 \times 1024$) directly in pixel-space. Building on the Transformer architecture, which is known to scale to billions of parameters, it bridges the gap between the efficiency of convolutional U-Nets and the scalability of Transformers. HDiT trains successfully without typical high-resolution training techniques such as multiscale architectures, latent autoencoders or self-conditioning. We demonstrate that HDiT performs competitively with existing models on ImageNet $256^2$, and sets a new state-of-the-art for diffusion models on FFHQ-$1024^2$.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11605
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers
Crowson, Katherine
Baumann, Stefan Andreas
Birch, Alex
Abraham, Tanishq Mathew
Kaplan, Daniel Z.
Shippole, Enrico
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We present the Hourglass Diffusion Transformer (HDiT), an image generative model that exhibits linear scaling with pixel count, supporting training at high-resolution (e.g. $1024 \times 1024$) directly in pixel-space. Building on the Transformer architecture, which is known to scale to billions of parameters, it bridges the gap between the efficiency of convolutional U-Nets and the scalability of Transformers. HDiT trains successfully without typical high-resolution training techniques such as multiscale architectures, latent autoencoders or self-conditioning. We demonstrate that HDiT performs competitively with existing models on ImageNet $256^2$, and sets a new state-of-the-art for diffusion models on FFHQ-$1024^2$.
title Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.11605