PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junsong, Ge, Chongjian, Xie, Enze, Wu, Yue, Yao, Lewei, Ren, Xiaozhe, Wang, Zhongdao, Luo, Ping, Lu, Huchuan, Li, Zhenguo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909138994331648
author Chen, Junsong
Ge, Chongjian
Xie, Enze
Wu, Yue
Yao, Lewei
Ren, Xiaozhe
Wang, Zhongdao
Luo, Ping
Lu, Huchuan
Li, Zhenguo
author_facet Chen, Junsong
Ge, Chongjian
Xie, Enze
Wu, Yue
Yao, Lewei
Ren, Xiaozhe
Wang, Zhongdao
Luo, Ping
Lu, Huchuan
Li, Zhenguo
contents In this paper, we introduce PixArt-Σ, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution. PixArt-Σrepresents a significant advancement over its predecessor, PixArt-α, offering images of markedly higher fidelity and improved alignment with text prompts. A key feature of PixArt-Σis its training efficiency. Leveraging the foundational pre-training of PixArt-α, it evolves from the `weaker' baseline to a `stronger' model via incorporating higher quality data, a process we term "weak-to-strong training". The advancements in PixArt-Σare twofold: (1) High-Quality Training Data: PixArt-Σincorporates superior-quality image data, paired with more precise and detailed image captions. (2) Efficient Token Compression: we propose a novel attention module within the DiT framework that compresses both keys and values, significantly improving efficiency and facilitating ultra-high-resolution image generation. Thanks to these improvements, PixArt-Σachieves superior image quality and user prompt adherence capabilities with significantly smaller model size (0.6B parameters) than existing text-to-image diffusion models, such as SDXL (2.6B parameters) and SD Cascade (5.1B parameters). Moreover, PixArt-Σ's capability to generate 4K images supports the creation of high-resolution posters and wallpapers, efficiently bolstering the production of high-quality visual content in industries such as film and gaming.
format Preprint
id arxiv_https___arxiv_org_abs_2403_04692
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Chen, Junsong
Ge, Chongjian
Xie, Enze
Wu, Yue
Yao, Lewei
Ren, Xiaozhe
Wang, Zhongdao
Luo, Ping
Lu, Huchuan
Li, Zhenguo
Computer Vision and Pattern Recognition
In this paper, we introduce PixArt-Σ, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution. PixArt-Σrepresents a significant advancement over its predecessor, PixArt-α, offering images of markedly higher fidelity and improved alignment with text prompts. A key feature of PixArt-Σis its training efficiency. Leveraging the foundational pre-training of PixArt-α, it evolves from the `weaker' baseline to a `stronger' model via incorporating higher quality data, a process we term "weak-to-strong training". The advancements in PixArt-Σare twofold: (1) High-Quality Training Data: PixArt-Σincorporates superior-quality image data, paired with more precise and detailed image captions. (2) Efficient Token Compression: we propose a novel attention module within the DiT framework that compresses both keys and values, significantly improving efficiency and facilitating ultra-high-resolution image generation. Thanks to these improvements, PixArt-Σachieves superior image quality and user prompt adherence capabilities with significantly smaller model size (0.6B parameters) than existing text-to-image diffusion models, such as SDXL (2.6B parameters) and SD Cascade (5.1B parameters). Moreover, PixArt-Σ's capability to generate 4K images supports the creation of high-resolution posters and wallpapers, efficiently bolstering the production of high-quality visual content in industries such as film and gaming.
title PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.04692