Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jungwon, Ko, Jungmin, Byun, Dongnam, Suh, Jangwon, Rhee, Wonjong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
by: Park, Jungwon, et al.
Published: (2026)
by: Park, Jungwon, et al.
Published: (2026)
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
by: Byun, Dongnam, et al.
Published: (2025)
by: Byun, Dongnam, et al.
Published: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
by: Kim, Jimyeong, et al.
Published: (2024)
by: Kim, Jimyeong, et al.
Published: (2024)
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
by: Kim, Jimyeong, et al.
Published: (2025)
by: Kim, Jimyeong, et al.
Published: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Evaluating Feature Attribution Methods for Electrocardiogram
by: Suh, Jangwon, et al.
Published: (2022)
by: Suh, Jangwon, et al.
Published: (2022)
Enhancing Contrastive Learning with Efficient Combinatorial Positive Pairing
by: Kim, Jaeill, et al.
Published: (2024)
by: Kim, Jaeill, et al.
Published: (2024)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
by: Song, Yeji, et al.
Published: (2024)
by: Song, Yeji, et al.
Published: (2024)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
by: Park, Jungwon, et al.
Published: (2026)
by: Park, Jungwon, et al.
Published: (2026)
On-Off Pattern Encoding and Path-Count Encoding as Deep Neural Network Representations
by: Jung, Euna, et al.
Published: (2024)
by: Jung, Euna, et al.
Published: (2024)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
by: Kim, Wonkyun, et al.
Published: (2024)
by: Kim, Wonkyun, et al.
Published: (2024)
Soft Head Selection for Injecting ICL-Derived Task Embeddings
by: Park, Jungwon, et al.
Published: (2025)
by: Park, Jungwon, et al.
Published: (2025)
Towards a Better Evaluation of Out-of-Domain Generalization
by: Hwang, Duhun, et al.
Published: (2024)
by: Hwang, Duhun, et al.
Published: (2024)
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
by: Kim, Jaeill, et al.
Published: (2024)
by: Kim, Jaeill, et al.
Published: (2024)
Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
by: Gao, Yunhe, et al.
Published: (2024)
by: Gao, Yunhe, et al.
Published: (2024)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)
by: Rahman, Tanzila, et al.
Published: (2024)
Unveiling Key Aspects of Fine-Tuning in Sentence Embeddings: A Representation Rank Analysis
by: Jung, Euna, et al.
Published: (2024)
by: Jung, Euna, et al.
Published: (2024)
TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
by: Koltsov, Kirill, et al.
Published: (2026)
by: Koltsov, Kirill, et al.
Published: (2026)
Multi-Head Attention Driven Dynamic Visual-Semantic Embedding for Enhanced Image-Text Matching
by: Chen, Wenjing
Published: (2024)
by: Chen, Wenjing
Published: (2024)
Instance-Aligned Captions for Explainable Video Anomaly Detection
by: Song, Inpyo, et al.
Published: (2026)
by: Song, Inpyo, et al.
Published: (2026)
NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior
by: Park, Dongwoo, et al.
Published: (2025)
by: Park, Dongwoo, et al.
Published: (2025)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Exploiting Text-Image Latent Spaces for the Description of Visual Concepts
by: Schmalwasser, Laines, et al.
Published: (2024)
by: Schmalwasser, Laines, et al.
Published: (2024)
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
by: Li, Ruihang, et al.
Published: (2026)
by: Li, Ruihang, et al.
Published: (2026)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
by: Zhang, Lei, et al.
Published: (2023)
by: Zhang, Lei, et al.
Published: (2023)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
PAFormer: Part Aware Transformer for Person Re-identification
by: Jung, Hyeono, et al.
Published: (2024)
by: Jung, Hyeono, et al.
Published: (2024)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
Anomaly Detection for People with Visual Impairments Using an Egocentric 360-Degree Camera
by: Song, Inpyo, et al.
Published: (2024)
by: Song, Inpyo, et al.
Published: (2024)
Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation
by: Tai, Ying, et al.
Published: (2025)
by: Tai, Ying, et al.
Published: (2025)
Personalized Residuals for Concept-Driven Text-to-Image Generation
by: Ham, Cusuh, et al.
Published: (2024)
by: Ham, Cusuh, et al.
Published: (2024)
Similar Items
-
Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
by: Park, Jungwon, et al.
Published: (2026) -
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
by: Byun, Dongnam, et al.
Published: (2025) -
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026) -
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
by: Kim, Jimyeong, et al.
Published: (2024) -
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
by: Kim, Jimyeong, et al.
Published: (2025)