PixelDiT: Pixel Diffusion Transformers for Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yongsheng, Xiong, Wei, Nie, Weili, Sheng, Yichen, Liu, Shiqiu, Luo, Jiebo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
DiP: Taming Diffusion Models in Pixel Space
by: Chen, Zhennan, et al.
Published: (2025)
by: Chen, Zhennan, et al.
Published: (2025)
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
by: Yu, Yongsheng, et al.
Published: (2024)
by: Yu, Yongsheng, et al.
Published: (2024)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
Registers Matter for Pixel-Space Diffusion Transformers
by: Starodubcev, Nikita, et al.
Published: (2026)
by: Starodubcev, Nikita, et al.
Published: (2026)
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion Models
by: Shih, Chun-Yen, et al.
Published: (2024)
by: Shih, Chun-Yen, et al.
Published: (2024)
Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
by: Zhou, Chenming, et al.
Published: (2025)
by: Zhou, Chenming, et al.
Published: (2025)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025)
by: Han, Minghao, et al.
Published: (2025)
Mapping Image Transformations Onto Pixel Processor Arrays
by: Bose, Laurie, et al.
Published: (2024)
by: Bose, Laurie, et al.
Published: (2024)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
by: Yang, Lianwei, et al.
Published: (2025)
by: Yang, Lianwei, et al.
Published: (2025)
On Inductive Biases That Enable Generalization of Diffusion Transformers
by: An, Jie, et al.
Published: (2024)
by: An, Jie, et al.
Published: (2024)
Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
by: Jamil, Sofia, et al.
Published: (2025)
by: Jamil, Sofia, et al.
Published: (2025)
FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
by: Lin, Mingfeng, et al.
Published: (2026)
by: Lin, Mingfeng, et al.
Published: (2026)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
by: Ma, Zehong, et al.
Published: (2026)
by: Ma, Zehong, et al.
Published: (2026)
Tracing Copied Pixels and Regularizing Patch Affinity in Copy Detection
by: Lu, Yichen, et al.
Published: (2026)
by: Lu, Yichen, et al.
Published: (2026)
PixelLM: Pixel Reasoning with Large Multimodal Model
by: Ren, Zhongwei, et al.
Published: (2023)
by: Ren, Zhongwei, et al.
Published: (2023)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
by: NVIDIA, et al.
Published: (2024)
by: NVIDIA, et al.
Published: (2024)
Dual-Scale Transformer for Large-Scale Single-Pixel Imaging
by: Qu, Gang, et al.
Published: (2024)
by: Qu, Gang, et al.
Published: (2024)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization
by: Guichemerre, Alexis, et al.
Published: (2025)
by: Guichemerre, Alexis, et al.
Published: (2025)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
by: Baade, Alan, et al.
Published: (2026)
by: Baade, Alan, et al.
Published: (2026)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
by: Jiang, Liyao, et al.
Published: (2024)
by: Jiang, Liyao, et al.
Published: (2024)
Aurora: Unified Video Editing with a Tool-Using Agent
by: Yu, Yongsheng, et al.
Published: (2026)
by: Yu, Yongsheng, et al.
Published: (2026)
PixNerd: Pixel Neural Field Diffusion
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
by: Sun, Haoze, et al.
Published: (2024)
by: Sun, Haoze, et al.
Published: (2024)
Transforming Weather Data from Pixel to Latent Space
by: Zhao, Sijie, et al.
Published: (2025)
by: Zhao, Sijie, et al.
Published: (2025)
PicoPose: Progressive Pixel-to-Pixel Correspondence Learning for Novel Object Pose Estimation
by: Liu, Lihua, et al.
Published: (2025)
by: Liu, Lihua, et al.
Published: (2025)
FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
Pixel-Aware Stable Diffusion for Realistic Image Super-resolution and Personalized Stylization
by: Yang, Tao, et al.
Published: (2023)
by: Yang, Tao, et al.
Published: (2023)
DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
by: Peng, Yichen, et al.
Published: (2026)
by: Peng, Yichen, et al.
Published: (2026)
Similar Items
-
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025) -
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
by: He, Yu, et al.
Published: (2026) -
DiP: Taming Diffusion Models in Pixel Space
by: Chen, Zhennan, et al.
Published: (2025) -
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025) -
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
by: Gong, Yan, et al.
Published: (2025)