PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junsong, Wu, Yue, Luo, Simian, Xie, Enze, Paul, Sayak, Luo, Ping, Zhao, Hang, Li, Zhenguo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913191549730816
author Chen, Junsong
Wu, Yue
Luo, Simian
Xie, Enze
Paul, Sayak
Luo, Ping
Zhao, Hang
Li, Zhenguo
author_facet Chen, Junsong
Wu, Yue
Luo, Simian
Xie, Enze
Paul, Sayak
Luo, Ping
Zhao, Hang
Li, Zhenguo
contents This technical report introduces PIXART-δ, a text-to-image synthesis framework that integrates the Latent Consistency Model (LCM) and ControlNet into the advanced PIXART-α model. PIXART-α is recognized for its ability to generate high-quality images of 1024px resolution through a remarkably efficient training process. The integration of LCM in PIXART-δ significantly accelerates the inference speed, enabling the production of high-quality images in just 2-4 steps. Notably, PIXART-δ achieves a breakthrough 0.5 seconds for generating 1024x1024 pixel images, marking a 7x improvement over the PIXART-α. Additionally, PIXART-δ is designed to be efficiently trainable on 32GB V100 GPUs within a single day. With its 8-bit inference capability (von Platen et al., 2023), PIXART-δ can synthesize 1024px images within 8GB GPU memory constraints, greatly enhancing its usability and accessibility. Furthermore, incorporating a ControlNet-like module enables fine-grained control over text-to-image diffusion models. We introduce a novel ControlNet-Transformer architecture, specifically tailored for Transformers, achieving explicit controllability alongside high-quality image generation. As a state-of-the-art, open-source image generation model, PIXART-δ offers a promising alternative to the Stable Diffusion family of models, contributing significantly to text-to-image synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05252
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
Chen, Junsong
Wu, Yue
Luo, Simian
Xie, Enze
Paul, Sayak
Luo, Ping
Zhao, Hang
Li, Zhenguo
Computer Vision and Pattern Recognition
This technical report introduces PIXART-δ, a text-to-image synthesis framework that integrates the Latent Consistency Model (LCM) and ControlNet into the advanced PIXART-α model. PIXART-α is recognized for its ability to generate high-quality images of 1024px resolution through a remarkably efficient training process. The integration of LCM in PIXART-δ significantly accelerates the inference speed, enabling the production of high-quality images in just 2-4 steps. Notably, PIXART-δ achieves a breakthrough 0.5 seconds for generating 1024x1024 pixel images, marking a 7x improvement over the PIXART-α. Additionally, PIXART-δ is designed to be efficiently trainable on 32GB V100 GPUs within a single day. With its 8-bit inference capability (von Platen et al., 2023), PIXART-δ can synthesize 1024px images within 8GB GPU memory constraints, greatly enhancing its usability and accessibility. Furthermore, incorporating a ControlNet-like module enables fine-grained control over text-to-image diffusion models. We introduce a novel ControlNet-Transformer architecture, specifically tailored for Transformers, achieving explicit controllability alongside high-quality image generation. As a state-of-the-art, open-source image generation model, PIXART-δ offers a promising alternative to the Stable Diffusion family of models, contributing significantly to text-to-image synthesis.
title PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.05252