PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junsong, Ge, Chongjian, Xie, Enze, Wu, Yue, Yao, Lewei, Ren, Xiaozhe, Wang, Zhongdao, Luo, Ping, Lu, Huchuan, Li, Zhenguo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023)
by: Chen, Junsong, et al.
Published: (2023)
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
by: Wang, Tianqi, et al.
Published: (2024)
by: Wang, Tianqi, et al.
Published: (2024)
Weak-to-Strong Diffusion with Reflection
by: Bai, Lichen, et al.
Published: (2025)
by: Bai, Lichen, et al.
Published: (2025)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
by: Xie, Enze, et al.
Published: (2024)
by: Xie, Enze, et al.
Published: (2024)
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
by: Zhao, Yuyang, et al.
Published: (2026)
by: Zhao, Yuyang, et al.
Published: (2026)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
by: Li, Yitong, et al.
Published: (2026)
by: Li, Yitong, et al.
Published: (2026)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023)
by: Huang, Kaiyi, et al.
Published: (2023)
Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
by: Nie, Fan, et al.
Published: (2025)
by: Nie, Fan, et al.
Published: (2025)
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
by: Zhu, Haoyi, et al.
Published: (2026)
by: Zhu, Haoyi, et al.
Published: (2026)
Quasi-aperiodic grain boundary phases of Σ5 tilt grain boundaries in refractory metals
by: Chen, Enze, et al.
Published: (2025)
by: Chen, Enze, et al.
Published: (2025)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
by: Hong, Susung, et al.
Published: (2025)
by: Hong, Susung, et al.
Published: (2025)
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
by: Xie, Enze, et al.
Published: (2025)
by: Xie, Enze, et al.
Published: (2025)
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
by: Chen, Junyu, et al.
Published: (2024)
by: Chen, Junyu, et al.
Published: (2024)
DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Accelerating Diffusion Sampling with Optimized Time Steps
by: Xue, Shuchen, et al.
Published: (2024)
by: Xue, Shuchen, et al.
Published: (2024)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
by: Tang, Haotian, et al.
Published: (2024)
by: Tang, Haotian, et al.
Published: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
by: Zhou, Gengze, et al.
Published: (2025)
by: Zhou, Gengze, et al.
Published: (2025)
LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
InstructRL4Pix: Training Diffusion for Image Editing by Reinforcement Learning
by: Li, Tiancheng, et al.
Published: (2024)
by: Li, Tiancheng, et al.
Published: (2024)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
Comprehensive analysis of the $γp \to K^+ Σ^0(1385)$, $γn \to K^+ Σ^-(1385)$, and $π^+ p \to K^+ Σ^+(1385)$ reactions
by: Wang, Ai-Chao, et al.
Published: (2025)
by: Wang, Ai-Chao, et al.
Published: (2025)
TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision
by: Gillani, Syeda Anshrah, et al.
Published: (2025)
by: Gillani, Syeda Anshrah, et al.
Published: (2025)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
Role of $Σ(1660)$ in the $K^- p \toπ^0π^0Σ^0$ reaction
by: Ji, Xing-Yi, et al.
Published: (2026)
by: Ji, Xing-Yi, et al.
Published: (2026)
Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
by: Li, Jianhao, et al.
Published: (2024)
by: Li, Jianhao, et al.
Published: (2024)
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
by: Xue, Shuchen, et al.
Published: (2025)
by: Xue, Shuchen, et al.
Published: (2025)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
by: He, Wenkun, et al.
Published: (2025)
by: He, Wenkun, et al.
Published: (2025)
PixNerd: Pixel Neural Field Diffusion
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
PixT3: Pixel-based Table-To-Text Generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
by: Panambur, Tejas, et al.
Published: (2025)
by: Panambur, Tejas, et al.
Published: (2025)
Investigation of Resonances in the $Σ({1/2}^{-})$ System Based on the Chiral Quark Model
by: Yao, Yu, et al.
Published: (2025)
by: Yao, Yu, et al.
Published: (2025)
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
by: Li, Xiaomin, et al.
Published: (2024)
by: Li, Xiaomin, et al.
Published: (2024)
Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation
by: Guo, Changlu, et al.
Published: (2025)
by: Guo, Changlu, et al.
Published: (2025)
AlgoFormer: An Efficient Transformer Framework with Algorithmic Structures
by: Gao, Yihang, et al.
Published: (2024)
by: Gao, Yihang, et al.
Published: (2024)
Similar Items
-
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023) -
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
by: Wang, Zhenyu, et al.
Published: (2024) -
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024) -
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024) -
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)