PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junsong, Yu, Jincheng, Ge, Chongjian, Yao, Lewei, Xie, Enze, Wu, Yue, Wang, Zhongdao, Kwok, James, Luo, Ping, Lu, Huchuan, Li, Zhenguo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
by: Zhao, Yuyang, et al.
Published: (2026)
by: Zhao, Yuyang, et al.
Published: (2026)
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
by: Zhu, Haoyi, et al.
Published: (2026)
by: Zhu, Haoyi, et al.
Published: (2026)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
by: Xie, Enze, et al.
Published: (2024)
by: Xie, Enze, et al.
Published: (2024)
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
by: Xie, Enze, et al.
Published: (2025)
by: Xie, Enze, et al.
Published: (2025)
SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
by: Wang, Tianqi, et al.
Published: (2024)
by: Wang, Tianqi, et al.
Published: (2024)
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
by: He, Wenkun, et al.
Published: (2025)
by: He, Wenkun, et al.
Published: (2025)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
by: Li, Yitong, et al.
Published: (2026)
by: Li, Yitong, et al.
Published: (2026)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023)
by: Huang, Kaiyi, et al.
Published: (2023)
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
by: Zhang, Kewei, et al.
Published: (2026)
by: Zhang, Kewei, et al.
Published: (2026)
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
by: Bai, Haoran, et al.
Published: (2025)
by: Bai, Haoran, et al.
Published: (2025)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
by: Hong, Susung, et al.
Published: (2025)
by: Hong, Susung, et al.
Published: (2025)
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
by: Chen, Junyu, et al.
Published: (2024)
by: Chen, Junyu, et al.
Published: (2024)
DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Accelerating Diffusion Sampling with Optimized Time Steps
by: Xue, Shuchen, et al.
Published: (2024)
by: Xue, Shuchen, et al.
Published: (2024)
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
by: Wu, Yecheng, et al.
Published: (2025)
by: Wu, Yecheng, et al.
Published: (2025)
MRI Scan Synthesis Methods based on Clustering and Pix2Pix
by: Baldini, Giulia, et al.
Published: (2023)
by: Baldini, Giulia, et al.
Published: (2023)
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
by: Wu, Chengyue, et al.
Published: (2026)
by: Wu, Chengyue, et al.
Published: (2026)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
by: Tang, Haotian, et al.
Published: (2024)
by: Tang, Haotian, et al.
Published: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
by: Zhou, Gengze, et al.
Published: (2025)
by: Zhou, Gengze, et al.
Published: (2025)
LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
InstructRL4Pix: Training Diffusion for Image Editing by Reinforcement Learning
by: Li, Tiancheng, et al.
Published: (2024)
by: Li, Tiancheng, et al.
Published: (2024)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision
by: Gillani, Syeda Anshrah, et al.
Published: (2025)
by: Gillani, Syeda Anshrah, et al.
Published: (2025)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
Fast-dLLM v2: Efficient Block-Diffusion LLM
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
by: Li, Jianhao, et al.
Published: (2024)
by: Li, Jianhao, et al.
Published: (2024)
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
by: Xue, Shuchen, et al.
Published: (2025)
by: Xue, Shuchen, et al.
Published: (2025)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
PixNerd: Pixel Neural Field Diffusion
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
PixT3: Pixel-based Table-To-Text Generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis
by: Teng, Yao, et al.
Published: (2024)
by: Teng, Yao, et al.
Published: (2024)
Similar Items
-
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024) -
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
by: Wang, Zhenyu, et al.
Published: (2024) -
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024) -
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024) -
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
by: Zhao, Yuyang, et al.
Published: (2026)