Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Bingyan, Wang, Chengyu, Cao, Tingfeng, Jia, Kui, Huang, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Attention Mechanism in Video Diffusion Models
von: Liu, Bingyan, et al.
Veröffentlicht: (2025)
von: Liu, Bingyan, et al.
Veröffentlicht: (2025)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
von: Zhan, Zechao, et al.
Veröffentlicht: (2024)
von: Zhan, Zechao, et al.
Veröffentlicht: (2024)
AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis
von: Cao, Shipeng, et al.
Veröffentlicht: (2026)
von: Cao, Shipeng, et al.
Veröffentlicht: (2026)
MIST: Mitigating Intersectional Bias with Disentangled Cross-Attention Editing in Text-to-Image Diffusion Models
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
von: Trippodo, Matteo, et al.
Veröffentlicht: (2025)
von: Trippodo, Matteo, et al.
Veröffentlicht: (2025)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
von: Wu, Shaojin, et al.
Veröffentlicht: (2024)
von: Wu, Shaojin, et al.
Veröffentlicht: (2024)
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
von: Soni, Achint, et al.
Veröffentlicht: (2025)
von: Soni, Achint, et al.
Veröffentlicht: (2025)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
von: Shuai, Xincheng, et al.
Veröffentlicht: (2024)
von: Shuai, Xincheng, et al.
Veröffentlicht: (2024)
Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image Editing
von: Koo, Gwanhyeong, et al.
Veröffentlicht: (2024)
von: Koo, Gwanhyeong, et al.
Veröffentlicht: (2024)
Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models
von: Jiang, Rui, et al.
Veröffentlicht: (2025)
von: Jiang, Rui, et al.
Veröffentlicht: (2025)
Forgedit: Text Guided Image Editing via Learning and Forgetting
von: Zhang, Shiwen, et al.
Veröffentlicht: (2023)
von: Zhang, Shiwen, et al.
Veröffentlicht: (2023)
AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2023)
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
von: Liang, Tianyi, et al.
Veröffentlicht: (2024)
von: Liang, Tianyi, et al.
Veröffentlicht: (2024)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
von: Ma, Lichen, et al.
Veröffentlicht: (2026)
von: Ma, Lichen, et al.
Veröffentlicht: (2026)
FastEdit: Fast Text-Guided Single-Image Editing via Semantic-Aware Diffusion Fine-Tuning
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects
von: Qiu, Weimin, et al.
Veröffentlicht: (2024)
von: Qiu, Weimin, et al.
Veröffentlicht: (2024)
FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing
von: Ren, Yufan, et al.
Veröffentlicht: (2025)
von: Ren, Yufan, et al.
Veröffentlicht: (2025)
Self-Attention Decomposition For Training Free Diffusion Editing
von: Anand, Tharun, et al.
Veröffentlicht: (2025)
von: Anand, Tharun, et al.
Veröffentlicht: (2025)
Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models
von: Xia, Tao, et al.
Veröffentlicht: (2026)
von: Xia, Tao, et al.
Veröffentlicht: (2026)
Exploring Text-Guided Single Image Editing for Remote Sensing Images
von: Han, Fangzhou, et al.
Veröffentlicht: (2024)
von: Han, Fangzhou, et al.
Veröffentlicht: (2024)
Text Guided Image Editing with Automatic Concept Locating and Forgetting
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance
von: Wang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Wang, Jiacheng, et al.
Veröffentlicht: (2024)
I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing
von: Yu, Jinghan, et al.
Veröffentlicht: (2026)
von: Yu, Jinghan, et al.
Veröffentlicht: (2026)
FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing
von: Yang, Kaixiang, et al.
Veröffentlicht: (2025)
von: Yang, Kaixiang, et al.
Veröffentlicht: (2025)
Guided Image Synthesis via Initial Image Editing in Diffusion Model
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
von: Guo, Xiefan, et al.
Veröffentlicht: (2026)
von: Guo, Xiefan, et al.
Veröffentlicht: (2026)
Diffusion Model-Based Image Editing: A Survey
von: Huang, Yi, et al.
Veröffentlicht: (2024)
von: Huang, Yi, et al.
Veröffentlicht: (2024)
UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
von: Bai, Chengyu, et al.
Veröffentlicht: (2025)
von: Bai, Chengyu, et al.
Veröffentlicht: (2025)
StableDrag: Stable Dragging for Point-based Image Editing
von: Cui, Yutao, et al.
Veröffentlicht: (2024)
von: Cui, Yutao, et al.
Veröffentlicht: (2024)
RSEdit: Text-Guided Image Editing for Remote Sensing
von: Zhenyuan, Chen, et al.
Veröffentlicht: (2026)
von: Zhenyuan, Chen, et al.
Veröffentlicht: (2026)
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
von: Liao, Wenhui, et al.
Veröffentlicht: (2024)
von: Liao, Wenhui, et al.
Veröffentlicht: (2024)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
Boosting Few-Shot Segmentation via Instance-Aware Data Augmentation and Local Consensus Guided Cross Attention
von: Guo, Li, et al.
Veröffentlicht: (2024)
von: Guo, Li, et al.
Veröffentlicht: (2024)
InstructGIE: Towards Generalizable Image Editing
von: Meng, Zichong, et al.
Veröffentlicht: (2024)
von: Meng, Zichong, et al.
Veröffentlicht: (2024)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
Group Relative Attention Guidance for Image Editing
von: Zhang, Xuanpu, et al.
Veröffentlicht: (2025)
von: Zhang, Xuanpu, et al.
Veröffentlicht: (2025)
Not All Regions Are Equal: Attention-Guided Perturbation Network for Industrial Anomaly Detection
von: Huang, Tingfeng, et al.
Veröffentlicht: (2024)
von: Huang, Tingfeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Understanding Attention Mechanism in Video Diffusion Models
von: Liu, Bingyan, et al.
Veröffentlicht: (2025) -
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024) -
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
von: Su, Tongtong, et al.
Veröffentlicht: (2025) -
MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
von: Zhan, Zechao, et al.
Veröffentlicht: (2024) -
AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis
von: Cao, Shipeng, et al.
Veröffentlicht: (2026)