DiP: Taming Diffusion Models in Pixel Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhennan, Zhu, Junwei, Chen, Xu, Zhang, Jiangning, Hu, Xiaobin, Zhao, Hanzhen, Wang, Chengjie, Yang, Jian, Tai, Ying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914423818420224
author Chen, Zhennan
Zhu, Junwei
Chen, Xu
Zhang, Jiangning
Hu, Xiaobin
Zhao, Hanzhen
Wang, Chengjie
Yang, Jian
Tai, Ying
author_facet Chen, Zhennan
Zhu, Junwei
Chen, Xu
Zhang, Jiangning
Hu, Xiaobin
Zhao, Hanzhen
Wang, Chengjie
Yang, Jian
Tai, Ying
contents Diffusion models face a fundamental trade-off between generation quality and computational efficiency. Latent Diffusion Models (LDMs) offer an efficient solution but suffer from potential information loss and non-end-to-end training. In contrast, existing pixel space models bypass VAEs but are computationally prohibitive for high-resolution synthesis. To resolve this dilemma, we propose DiP, an efficient pixel space diffusion framework. DiP decouples generation into a global and a local stage: a Diffusion Transformer (DiT) backbone operates on large patches for efficient global structure construction, while a co-trained lightweight Patch Detailer Head leverages contextual features to restore fine-grained local details. This synergistic design achieves computational efficiency comparable to LDMs without relying on a VAE. DiP is accomplished with up to 10$\times$ faster inference speeds than previous method while increasing the total number of parameters by only 0.3%, and achieves an 1.79 FID score on ImageNet 256$\times$256.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiP: Taming Diffusion Models in Pixel Space
Chen, Zhennan
Zhu, Junwei
Chen, Xu
Zhang, Jiangning
Hu, Xiaobin
Zhao, Hanzhen
Wang, Chengjie
Yang, Jian
Tai, Ying
Computer Vision and Pattern Recognition
Diffusion models face a fundamental trade-off between generation quality and computational efficiency. Latent Diffusion Models (LDMs) offer an efficient solution but suffer from potential information loss and non-end-to-end training. In contrast, existing pixel space models bypass VAEs but are computationally prohibitive for high-resolution synthesis. To resolve this dilemma, we propose DiP, an efficient pixel space diffusion framework. DiP decouples generation into a global and a local stage: a Diffusion Transformer (DiT) backbone operates on large patches for efficient global structure construction, while a co-trained lightweight Patch Detailer Head leverages contextual features to restore fine-grained local details. This synergistic design achieves computational efficiency comparable to LDMs without relying on a VAE. DiP is accomplished with up to 10$\times$ faster inference speeds than previous method while increasing the total number of parameters by only 0.3%, and achieves an 1.79 FID score on ImageNet 256$\times$256.
title DiP: Taming Diffusion Models in Pixel Space
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18822