Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jeeyung, Esmaeili, Erfan, Qiu, Qiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Constructing Concept-based Models to Mitigate Spurious Correlations with Minimal Human Effort
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
BiasMap: Leveraging Cross-Attentions to Discover and Mitigate Hidden Social Biases in Text-to-Image Generation
by: Chakraborty, Rajatsubhra, et al.
Published: (2025)
by: Chakraborty, Rajatsubhra, et al.
Published: (2025)
Information Theoretic Text-to-Image Alignment
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting
by: Colombo, Antonio, et al.
Published: (2026)
by: Colombo, Antonio, et al.
Published: (2026)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
TempoControl: Temporal Attention Guidance for Text-to-Video Models
by: Schiber, Shira, et al.
Published: (2025)
by: Schiber, Shira, et al.
Published: (2025)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
by: Zhang, Yunzhe, et al.
Published: (2026)
by: Zhang, Yunzhe, et al.
Published: (2026)
Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
by: Dahary, Omer, et al.
Published: (2024)
by: Dahary, Omer, et al.
Published: (2024)
DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization
by: Nam, Jisu, et al.
Published: (2024)
by: Nam, Jisu, et al.
Published: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Aligning Text to Image in Diffusion Models is Easier Than You Think
by: Lee, Jaa-Yeon, et al.
Published: (2025)
by: Lee, Jaa-Yeon, et al.
Published: (2025)
RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
by: Izadi, Amir Mohammad, et al.
Published: (2025)
by: Izadi, Amir Mohammad, et al.
Published: (2025)
Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
by: Lee, Byung Hyun, et al.
Published: (2025)
by: Lee, Byung Hyun, et al.
Published: (2025)
Expressive Text-to-Image Generation with Rich Text
by: Ge, Songwei, et al.
Published: (2023)
by: Ge, Songwei, et al.
Published: (2023)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Semimage: HSV-Based Semantic Image Encoding for Disentangled Text Representation
by: Zare, Mohammad
Published: (2025)
by: Zare, Mohammad
Published: (2025)
TextCraftor: Your Text Encoder Can be Image Quality Controller
by: Li, Yanyu, et al.
Published: (2024)
by: Li, Yanyu, et al.
Published: (2024)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
TextCAM: Explaining Class Activation Map with Text
by: Zhao, Qiming, et al.
Published: (2025)
by: Zhao, Qiming, et al.
Published: (2025)
Model-Agnostic Human Preference Inversion in Diffusion Models
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
by: Bahng, Hyojin, et al.
Published: (2025)
by: Bahng, Hyojin, et al.
Published: (2025)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024)
by: Zhang, Dinghuai, et al.
Published: (2024)
Text-Guided Image Clustering
by: Stephan, Andreas, et al.
Published: (2024)
by: Stephan, Andreas, et al.
Published: (2024)
Hyperbolic Image-Text Representations
by: Desai, Karan, et al.
Published: (2023)
by: Desai, Karan, et al.
Published: (2023)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
by: Lee, Jaa-Yeon, et al.
Published: (2026)
by: Lee, Jaa-Yeon, et al.
Published: (2026)
Controlling Text-to-Image Diffusion by Orthogonal Finetuning
by: Qiu, Zeju, et al.
Published: (2023)
by: Qiu, Zeju, et al.
Published: (2023)
Directional Textual Inversion for Personalized Text-to-Image Generation
by: Kim, Kunhee, et al.
Published: (2025)
by: Kim, Kunhee, et al.
Published: (2025)
ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
by: Xing, Eric, et al.
Published: (2025)
by: Xing, Eric, et al.
Published: (2025)
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
by: Ganjdanesh, Alireza, et al.
Published: (2024)
by: Ganjdanesh, Alireza, et al.
Published: (2024)
All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
by: Li, Shuangqi, et al.
Published: (2024)
by: Li, Shuangqi, et al.
Published: (2024)
Learning to Rank Caption Chains for Video-Text Alignment
by: Blume, Ansel, et al.
Published: (2026)
by: Blume, Ansel, et al.
Published: (2026)
Similar Items
-
Constructing Concept-based Models to Mitigate Spurious Correlations with Minimal Human Effort
by: Kim, Jeeyung, et al.
Published: (2024) -
Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
by: Kim, Taehoon, et al.
Published: (2025) -
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024) -
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025) -
BiasMap: Leveraging Cross-Attentions to Discover and Mitigate Hidden Social Biases in Text-to-Image Generation
by: Chakraborty, Rajatsubhra, et al.
Published: (2025)