PixelDiT: Pixel Diffusion Transformers for Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Yongsheng, Xiong, Wei, Nie, Weili, Sheng, Yichen, Liu, Shiqiu, Luo, Jiebo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915939955507200
author Yu, Yongsheng
Xiong, Wei
Nie, Weili
Sheng, Yichen
Liu, Shiqiu
Luo, Jiebo
author_facet Yu, Yongsheng
Xiong, Wei
Nie, Weili
Sheng, Yichen
Liu, Shiqiu
Luo, Jiebo
contents Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoencoder introduces lossy reconstruction, leading to error accumulation while hindering joint optimization. To address these issues, we propose PixelDiT, a single-stage, end-to-end model that eliminates the need for the autoencoder and learns the diffusion process directly in the pixel space. PixelDiT adopts a fully transformer-based architecture shaped by a dual-level design: a patch-level DiT that captures global semantics and a pixel-level DiT that refines texture details, enabling efficient training of a pixel-space diffusion model while preserving fine details. PixelDiT achieves 1.61 FID on ImageNet 256 and 1.81 FID on ImageNet 512, surpassing existing pixel generative models. We further extend PixelDiT to text-to-image generation and pretrain it at the 10242resolution in pixel space. It achieves 0.74 on GenEval and 83.5 on DPG-bench, approaching the best latent diffusion models. Code: https://github.com/NVlabs/PixelDiT
format Preprint
id arxiv_https___arxiv_org_abs_2511_20645
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PixelDiT: Pixel Diffusion Transformers for Image Generation
Yu, Yongsheng
Xiong, Wei
Nie, Weili
Sheng, Yichen
Liu, Shiqiu
Luo, Jiebo
Computer Vision and Pattern Recognition
Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoencoder introduces lossy reconstruction, leading to error accumulation while hindering joint optimization. To address these issues, we propose PixelDiT, a single-stage, end-to-end model that eliminates the need for the autoencoder and learns the diffusion process directly in the pixel space. PixelDiT adopts a fully transformer-based architecture shaped by a dual-level design: a patch-level DiT that captures global semantics and a pixel-level DiT that refines texture details, enabling efficient training of a pixel-space diffusion model while preserving fine details. PixelDiT achieves 1.61 FID on ImageNet 256 and 1.81 FID on ImageNet 512, surpassing existing pixel generative models. We further extend PixelDiT to text-to-image generation and pretrain it at the 10242resolution in pixel space. It achieves 0.74 on GenEval and 83.5 on DPG-bench, approaching the best latent diffusion models. Code: https://github.com/NVlabs/PixelDiT
title PixelDiT: Pixel Diffusion Transformers for Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20645