HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yu, Tang, Fan, Cao, Juan, Zhang, Yuxin, Kong, Xiaoyu, Li, Jintao, Deussen, Oliver, Lee, Tong-Yee
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909399948197888
author Xu, Yu
Tang, Fan
Cao, Juan
Zhang, Yuxin
Kong, Xiaoyu
Li, Jintao
Deussen, Oliver
Lee, Tong-Yee
author_facet Xu, Yu
Tang, Fan
Cao, Juan
Zhang, Yuxin
Kong, Xiaoyu
Li, Jintao
Deussen, Oliver
Lee, Tong-Yee
contents Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter, a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we present a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experimental results on multiple benchmarks demonstrate HeadRouter's performance in terms of editing fidelity and image quality.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15034
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
Xu, Yu
Tang, Fan
Cao, Juan
Zhang, Yuxin
Kong, Xiaoyu
Li, Jintao
Deussen, Oliver
Lee, Tong-Yee
Computer Vision and Pattern Recognition
Machine Learning
Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter, a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we present a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experimental results on multiple benchmarks demonstrate HeadRouter's performance in terms of editing fidelity and image quality.
title HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.15034