LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lian, Long, Li, Boyi, Yala, Adam, Darrell, Trevor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
Rethinking Patch Dependence for Masked Autoencoders
von: Fu, Letian, et al.
Veröffentlicht: (2024)
von: Fu, Letian, et al.
Veröffentlicht: (2024)
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
von: Zhang, Gaoyang, et al.
Veröffentlicht: (2024)
von: Zhang, Gaoyang, et al.
Veröffentlicht: (2024)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
von: Wu, Chen, et al.
Veröffentlicht: (2024)
von: Wu, Chen, et al.
Veröffentlicht: (2024)
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
von: Fan, Zezhong, et al.
Veröffentlicht: (2024)
von: Fan, Zezhong, et al.
Veröffentlicht: (2024)
EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
von: Ma, Bingqi, et al.
Veröffentlicht: (2024)
von: Ma, Bingqi, et al.
Veröffentlicht: (2024)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
von: Nguyen, Viet, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet, et al.
Veröffentlicht: (2024)
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
von: Li, Hongyu, et al.
Veröffentlicht: (2024)
von: Li, Hongyu, et al.
Veröffentlicht: (2024)
Discovering Divergent Representations between Text-to-Image Models
von: Dunlap, Lisa, et al.
Veröffentlicht: (2025)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2025)
Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model
von: Yi, Mingyang, et al.
Veröffentlicht: (2024)
von: Yi, Mingyang, et al.
Veröffentlicht: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
Image Super-Resolution with Text Prompt Diffusion
von: Chen, Zheng, et al.
Veröffentlicht: (2023)
von: Chen, Zheng, et al.
Veröffentlicht: (2023)
LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2025)
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2025)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
von: Tang, Bingda, et al.
Veröffentlicht: (2025)
von: Tang, Bingda, et al.
Veröffentlicht: (2025)
DreamDrone: Text-to-Image Diffusion Models are Zero-shot Perpetual View Generators
von: Kong, Hanyang, et al.
Veröffentlicht: (2023)
von: Kong, Hanyang, et al.
Veröffentlicht: (2023)
Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
Debiasing Text-to-Image Diffusion Models
von: He, Ruifei, et al.
Veröffentlicht: (2024)
von: He, Ruifei, et al.
Veröffentlicht: (2024)
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
von: Chen, Xinyan, et al.
Veröffentlicht: (2023)
von: Chen, Xinyan, et al.
Veröffentlicht: (2023)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
von: Chin, Zhi-Yi, et al.
Veröffentlicht: (2023)
von: Chin, Zhi-Yi, et al.
Veröffentlicht: (2023)
Equipping Diffusion Models with Differentiable Spatial Entropy for Low-Light Image Enhancement
von: Lian, Wenyi, et al.
Veröffentlicht: (2024)
von: Lian, Wenyi, et al.
Veröffentlicht: (2024)
Wuerstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
von: Pernias, Pablo, et al.
Veröffentlicht: (2023)
von: Pernias, Pablo, et al.
Veröffentlicht: (2023)
ECNet: Effective Controllable Text-to-Image Diffusion Models
von: Li, Sicheng, et al.
Veröffentlicht: (2024)
von: Li, Sicheng, et al.
Veröffentlicht: (2024)
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
von: Wang, Zhendong, et al.
Veröffentlicht: (2025)
von: Wang, Zhendong, et al.
Veröffentlicht: (2025)
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
von: Ganjdanesh, Alireza, et al.
Veröffentlicht: (2024)
von: Ganjdanesh, Alireza, et al.
Veröffentlicht: (2024)
Readout Guidance: Learning Control from Diffusion Features
von: Luo, Grace, et al.
Veröffentlicht: (2023)
von: Luo, Grace, et al.
Veröffentlicht: (2023)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models
von: Manukyan, Hayk, et al.
Veröffentlicht: (2023)
von: Manukyan, Hayk, et al.
Veröffentlicht: (2023)
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
von: Han, Yucheng, et al.
Veröffentlicht: (2024)
von: Han, Yucheng, et al.
Veröffentlicht: (2024)
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
von: Xu, Xingqian, et al.
Veröffentlicht: (2022)
von: Xu, Xingqian, et al.
Veröffentlicht: (2022)
Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
von: Li, Ruibin, et al.
Veröffentlicht: (2024)
von: Li, Ruibin, et al.
Veröffentlicht: (2024)
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
von: Li, Shufan, et al.
Veröffentlicht: (2025)
von: Li, Shufan, et al.
Veröffentlicht: (2025)
Blending Concepts with Text-to-Image Diffusion Models
von: Olearo, Lorenzo, et al.
Veröffentlicht: (2025)
von: Olearo, Lorenzo, et al.
Veröffentlicht: (2025)
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
von: Luo, Grace, et al.
Veröffentlicht: (2023)
von: Luo, Grace, et al.
Veröffentlicht: (2023)
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
von: Harrington, Anne, et al.
Veröffentlicht: (2025)
von: Harrington, Anne, et al.
Veröffentlicht: (2025)
Pillar-0: A New Frontier for Radiology Foundation Models
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023) -
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025) -
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025) -
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025) -
Rethinking Patch Dependence for Masked Autoencoders
von: Fu, Letian, et al.
Veröffentlicht: (2024)