Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hao, Lal, Shamit, Li, Zhiheng, Xie, Yusheng, Wang, Ying, Zou, Yang, Majumder, Orchid, Manmatha, R., Tu, Zhuowen, Ermon, Stefano, Soatto, Stefano, Swaminathan, Ashwin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Mixed-Query Transformer: A Unified Image Segmentation Architecture
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024)
by: Li, Xiaolong, et al.
Published: (2024)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023)
by: Golatkar, Aditya, et al.
Published: (2023)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
by: Shi, Kunyu, et al.
Published: (2024)
by: Shi, Kunyu, et al.
Published: (2024)
A Quantitative Evaluation of Score Distillation Sampling Based Text-to-3D
by: Fei, Xiaohan, et al.
Published: (2024)
by: Fei, Xiaohan, et al.
Published: (2024)
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
by: Tao, Chaofan, et al.
Published: (2024)
by: Tao, Chaofan, et al.
Published: (2024)
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
by: Tan, Jing, et al.
Published: (2026)
by: Tan, Jing, et al.
Published: (2026)
Fast Sparse View Guided NeRF Update for Object Reconfigurations
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)
by: Golatkar, Aditya, et al.
Published: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)
by: Ram, Dhananjay, et al.
Published: (2025)
RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
by: Chen, Tianlang, et al.
Published: (2025)
by: Chen, Tianlang, et al.
Published: (2025)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
by: Becker, Evan, et al.
Published: (2024)
by: Becker, Evan, et al.
Published: (2024)
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
by: Lou, Aaron, et al.
Published: (2023)
by: Lou, Aaron, et al.
Published: (2023)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
Energy Scaling Laws for Diffusion Models: Quantifying Compute in Image Generation
by: Iyengar, Aniketh, et al.
Published: (2025)
by: Iyengar, Aniketh, et al.
Published: (2025)
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
by: Zhao, Yue, et al.
Published: (2022)
by: Zhao, Yue, et al.
Published: (2022)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
by: Guo, Gabe, et al.
Published: (2025)
by: Guo, Gabe, et al.
Published: (2025)
ExPLoRA: Parameter-Efficient Extended Pre-Training to Adapt Vision Transformers under Domain Shifts
by: Khanna, Samar, et al.
Published: (2024)
by: Khanna, Samar, et al.
Published: (2024)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
by: Wang, Zirui, et al.
Published: (2023)
by: Wang, Zirui, et al.
Published: (2023)
Divergence Minimization Preference Optimization for Diffusion Model Alignment
by: Li, Binxu, et al.
Published: (2025)
by: Li, Binxu, et al.
Published: (2025)
DreamPropeller: Supercharge Text-to-3D Generation with Parallel Sampling
by: Zhou, Linqi, et al.
Published: (2023)
by: Zhou, Linqi, et al.
Published: (2023)
Deep Learning for MRI Slice Interpolation: The Critical Role of Problem Formulation
by: Savant, Shamit
Published: (2026)
by: Savant, Shamit
Published: (2026)
FairRAG: Fair Human Generation via Fair Retrieval Augmentation
by: Shrestha, Robik, et al.
Published: (2024)
by: Shrestha, Robik, et al.
Published: (2024)
Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Energy-Based Diffusion Language Models for Text Generation
by: Xu, Minkai, et al.
Published: (2024)
by: Xu, Minkai, et al.
Published: (2024)
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
by: Wang, Austin, et al.
Published: (2026)
by: Wang, Austin, et al.
Published: (2026)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Exploring Human Quadruped Locomotion for Exergames
by: Ahmed, Shamit, et al.
Published: (2026)
by: Ahmed, Shamit, et al.
Published: (2026)
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
by: Srivastava, Divyansh, et al.
Published: (2025)
by: Srivastava, Divyansh, et al.
Published: (2025)
Geometric Trajectory Diffusion Models
by: Han, Jiaqi, et al.
Published: (2024)
by: Han, Jiaqi, et al.
Published: (2024)
Adaptive Spectral Feature Forecasting for Diffusion Sampling Acceleration
by: Han, Jiaqi, et al.
Published: (2026)
by: Han, Jiaqi, et al.
Published: (2026)
Personalized Preference Fine-tuning of Diffusion Models
by: Dang, Meihua, et al.
Published: (2025)
by: Dang, Meihua, et al.
Published: (2025)
TrAct: Making First-layer Pre-Activations Trainable
by: Petersen, Felix, et al.
Published: (2024)
by: Petersen, Felix, et al.
Published: (2024)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Similar Items
-
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024) -
Mixed-Query Transformer: A Unified Image Segmentation Architecture
by: Wang, Pei, et al.
Published: (2024) -
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024) -
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024) -
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023)