DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Jiatao, Wang, Yuyang, Zhang, Yizhe, Zhang, Qihang, Zhang, Dinghuai, Jaitly, Navdeep, Susskind, Josh, Zhai, Shuangfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving GFlowNets for Text-to-Image Diffusion Alignment
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024)
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024)
Matryoshka Diffusion Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2023)
von: Gu, Jiatao, et al.
Veröffentlicht: (2023)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024)
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
How Far Are We from Intelligent Visual Deductive Reasoning?
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
Normalizing Flows are Capable Generative Models
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024)
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024)
Many-to-many Image Generation with Auto-regressive Diffusion Models
von: Shen, Ying, et al.
Veröffentlicht: (2024)
von: Shen, Ying, et al.
Veröffentlicht: (2024)
Normalizing Trajectory Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2026)
von: Gu, Jiatao, et al.
Veröffentlicht: (2026)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
von: Shen, Ying, et al.
Veröffentlicht: (2026)
von: Shen, Ying, et al.
Veröffentlicht: (2026)
Normalizing Flows with Iterative Denoising
von: Chen, Tianrong, et al.
Veröffentlicht: (2026)
von: Chen, Tianrong, et al.
Veröffentlicht: (2026)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
PLANNER: Generating Diversified Paragraph via Latent Language Diffusion Model
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2023)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
World-consistent Video Diffusion with Explicit 3D Modeling
von: Zhang, Qihang, et al.
Veröffentlicht: (2024)
von: Zhang, Qihang, et al.
Veröffentlicht: (2024)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
von: Berthelot, David, et al.
Veröffentlicht: (2026)
von: Berthelot, David, et al.
Veröffentlicht: (2026)
Scalable Pre-training of Large Autoregressive Image Models
von: El-Nouby, Alaaeldin, et al.
Veröffentlicht: (2024)
von: El-Nouby, Alaaeldin, et al.
Veröffentlicht: (2024)
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models
von: Miao, Kevin, et al.
Veröffentlicht: (2024)
von: Miao, Kevin, et al.
Veröffentlicht: (2024)
Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
Learning to Sample Effective and Diverse Prompts for Text-to-Image Generation
von: Yun, Taeyoung, et al.
Veröffentlicht: (2025)
von: Yun, Taeyoung, et al.
Veröffentlicht: (2025)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
von: Huang, Chen, et al.
Veröffentlicht: (2026)
von: Huang, Chen, et al.
Veröffentlicht: (2026)
MixAR: Mixture Autoregressive Image Generation
von: Hu, Jinyuan, et al.
Veröffentlicht: (2025)
von: Hu, Jinyuan, et al.
Veröffentlicht: (2025)
3D Shape Tokenization via Latent Flow Matching
von: Chang, Jen-Hao Rick, et al.
Veröffentlicht: (2024)
von: Chang, Jen-Hao Rick, et al.
Veröffentlicht: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
Autoregressive Image Generation with Masked Bit Modeling
von: Yu, Qihang, et al.
Veröffentlicht: (2026)
von: Yu, Qihang, et al.
Veröffentlicht: (2026)
Denoising Autoregressive Representation Learning
von: Li, Yazhe, et al.
Veröffentlicht: (2024)
von: Li, Yazhe, et al.
Veröffentlicht: (2024)
Xformer: Hybrid X-Shaped Transformer for Image Denoising
von: Zhang, Jiale, et al.
Veröffentlicht: (2023)
von: Zhang, Jiale, et al.
Veröffentlicht: (2023)
MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2025)
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2025)
Astra: General Interactive World Model with Autoregressive Denoising
von: Zhu, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yixuan, et al.
Veröffentlicht: (2025)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
von: Sun, Peize, et al.
Veröffentlicht: (2024)
von: Sun, Peize, et al.
Veröffentlicht: (2024)
Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards
von: Liu, Qingming, et al.
Veröffentlicht: (2025)
von: Liu, Qingming, et al.
Veröffentlicht: (2025)
Scalable Autoregressive Image Generation with Mamba
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
von: Teng, Yao, et al.
Veröffentlicht: (2025)
von: Teng, Yao, et al.
Veröffentlicht: (2025)
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
von: Chu, Wenda, et al.
Veröffentlicht: (2026)
von: Chu, Wenda, et al.
Veröffentlicht: (2026)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
von: Tschannen, Michael, et al.
Veröffentlicht: (2024)
von: Tschannen, Michael, et al.
Veröffentlicht: (2024)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
TADA: Improved Diffusion Sampling with Training-free Augmented Dynamics
von: Chen, Tianrong, et al.
Veröffentlicht: (2025)
von: Chen, Tianrong, et al.
Veröffentlicht: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Improving GFlowNets for Text-to-Image Diffusion Alignment
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024) -
Matryoshka Diffusion Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2023) -
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024) -
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
von: Gu, Jiatao, et al.
Veröffentlicht: (2024) -
How Far Are We from Intelligent Visual Deductive Reasoning?
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)