Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Bingda, Zheng, Boyang, Pan, Xichen, Paul, Sayak, Xie, Saining |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
di: Tong, Shengbang, et al.
Pubblicazione: (2026)
di: Tong, Shengbang, et al.
Pubblicazione: (2026)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
di: Li, Chenyu, et al.
Pubblicazione: (2025)
di: Li, Chenyu, et al.
Pubblicazione: (2025)
Diffusion Transformers with Representation Autoencoders
di: Zheng, Boyang, et al.
Pubblicazione: (2025)
di: Zheng, Boyang, et al.
Pubblicazione: (2025)
Image Sculpting: Precise Object Editing with 3D Geometry Control
di: Yenphraphai, Jiraphon, et al.
Pubblicazione: (2024)
di: Yenphraphai, Jiraphon, et al.
Pubblicazione: (2024)
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
di: Ma, Nanye, et al.
Pubblicazione: (2024)
di: Ma, Nanye, et al.
Pubblicazione: (2024)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
di: Pan, Xichen, et al.
Pubblicazione: (2023)
di: Pan, Xichen, et al.
Pubblicazione: (2023)
SARD: Segmentation-Aware Anomaly Synthesis via Region-Constrained Diffusion with Discriminative Mask Guidance
di: Wang, Yanshu, et al.
Pubblicazione: (2025)
di: Wang, Yanshu, et al.
Pubblicazione: (2025)
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
di: Luo, Lingxiao, et al.
Pubblicazione: (2024)
di: Luo, Lingxiao, et al.
Pubblicazione: (2024)
Cambrian-P: Pose-Grounded Video Understanding
di: Yang, Jihan, et al.
Pubblicazione: (2026)
di: Yang, Jihan, et al.
Pubblicazione: (2026)
LDGen: Enhancing Text-to-Image Synthesis via Large Language Model-Driven Language Representation
di: Li, Pengzhi, et al.
Pubblicazione: (2025)
di: Li, Pengzhi, et al.
Pubblicazione: (2025)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
di: Leng, Xingjian, et al.
Pubblicazione: (2025)
di: Leng, Xingjian, et al.
Pubblicazione: (2025)
DiffusionGuard: A Robust Defense Against Malicious Diffusion-based Image Editing
di: Choi, June Suk, et al.
Pubblicazione: (2024)
di: Choi, June Suk, et al.
Pubblicazione: (2024)
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
di: Chen, Junsong, et al.
Pubblicazione: (2023)
di: Chen, Junsong, et al.
Pubblicazione: (2023)
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
di: Zhuo, Le, et al.
Pubblicazione: (2025)
di: Zhuo, Le, et al.
Pubblicazione: (2025)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
di: Chen, Xinlei, et al.
Pubblicazione: (2024)
di: Chen, Xinlei, et al.
Pubblicazione: (2024)
LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
di: Dong, Xuanzhao, et al.
Pubblicazione: (2025)
di: Dong, Xuanzhao, et al.
Pubblicazione: (2025)
Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects
di: Qiu, Weimin, et al.
Pubblicazione: (2024)
di: Qiu, Weimin, et al.
Pubblicazione: (2024)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
di: Yang, Jihan, et al.
Pubblicazione: (2024)
di: Yang, Jihan, et al.
Pubblicazione: (2024)
Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
di: Zheng, Youwei, et al.
Pubblicazione: (2025)
di: Zheng, Youwei, et al.
Pubblicazione: (2025)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
di: Lian, Long, et al.
Pubblicazione: (2023)
di: Lian, Long, et al.
Pubblicazione: (2023)
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
di: Xie, Chang, et al.
Pubblicazione: (2025)
di: Xie, Chang, et al.
Pubblicazione: (2025)
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
di: Ding, Xichen, et al.
Pubblicazione: (2025)
di: Ding, Xichen, et al.
Pubblicazione: (2025)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
di: Xie, Enze, et al.
Pubblicazione: (2024)
di: Xie, Enze, et al.
Pubblicazione: (2024)
Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment
di: Xie, Xing, et al.
Pubblicazione: (2025)
di: Xie, Xing, et al.
Pubblicazione: (2025)
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
di: Liu, Bingchen, et al.
Pubblicazione: (2024)
di: Liu, Bingchen, et al.
Pubblicazione: (2024)
Building Universal Foundation Models for Medical Image Analysis with Spatially Adaptive Networks
di: Luo, Lingxiao, et al.
Pubblicazione: (2023)
di: Luo, Lingxiao, et al.
Pubblicazione: (2023)
Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic Segmentation
di: Xia, Ruihao, et al.
Pubblicazione: (2024)
di: Xia, Ruihao, et al.
Pubblicazione: (2024)
FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis
di: Xu, Xichen, et al.
Pubblicazione: (2025)
di: Xu, Xichen, et al.
Pubblicazione: (2025)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
di: Zhangli, Qilong, et al.
Pubblicazione: (2024)
di: Zhangli, Qilong, et al.
Pubblicazione: (2024)
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
di: Chen, Jiacheng, et al.
Pubblicazione: (2025)
di: Chen, Jiacheng, et al.
Pubblicazione: (2025)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
DreamFuse: Adaptive Image Fusion with Diffusion Transformer
di: Huang, Junjia, et al.
Pubblicazione: (2025)
di: Huang, Junjia, et al.
Pubblicazione: (2025)
LM4LV: A Frozen Large Language Model for Low-level Vision Tasks
di: Zheng, Boyang, et al.
Pubblicazione: (2024)
di: Zheng, Boyang, et al.
Pubblicazione: (2024)
LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model
di: Huang, Runhui, et al.
Pubblicazione: (2024)
di: Huang, Runhui, et al.
Pubblicazione: (2024)
MaxFusion: Plug&Play Multi-Modal Generation in Text-to-Image Diffusion Models
di: Nair, Nithin Gopalakrishnan, et al.
Pubblicazione: (2024)
di: Nair, Nithin Gopalakrishnan, et al.
Pubblicazione: (2024)
Test-time Conditional Text-to-Image Synthesis Using Diffusion Models
di: Shukla, Tripti, et al.
Pubblicazione: (2024)
di: Shukla, Tripti, et al.
Pubblicazione: (2024)
Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
di: Kim, Jin Hyeon, et al.
Pubblicazione: (2025)
di: Kim, Jin Hyeon, et al.
Pubblicazione: (2025)
Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model
di: Zhang, Hao, et al.
Pubblicazione: (2024)
di: Zhang, Hao, et al.
Pubblicazione: (2024)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
Exploring State Space Model in Wavelet Domain: An Infrared and Visible Image Fusion Network via Wavelet Transform and State Space Model
di: Zhang, Tianpei, et al.
Pubblicazione: (2025)
di: Zhang, Tianpei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
di: Tong, Shengbang, et al.
Pubblicazione: (2026) -
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
di: Li, Chenyu, et al.
Pubblicazione: (2025) -
Diffusion Transformers with Representation Autoencoders
di: Zheng, Boyang, et al.
Pubblicazione: (2025) -
Image Sculpting: Precise Object Editing with 3D Geometry Control
di: Yenphraphai, Jiraphon, et al.
Pubblicazione: (2024) -
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
di: Ma, Nanye, et al.
Pubblicazione: (2024)