Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.21090 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912558014791680 |
|---|---|
| author | Kim, Namu Kweon, Wonbin Kim, Minsoo Yu, Hwanjo |
| author_facet | Kim, Namu Kweon, Wonbin Kim, Minsoo Yu, Hwanjo |
| contents | We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by the Query-Key alignment. To tackle this issue, we introduce Q-Align, utilizing Query-Query alignment to mitigate attention leakage and improve the semantic alignment in zero-shot appearance transfer. Q-Align incorporates three core contributions: (1) Query-Query alignment, facilitating the sophisticated spatial semantic mapping between two images; (2) Key-Value rearrangement, enhancing feature correspondence through realignment; and (3) Attention refinement using rearranged keys and values to maintain semantic consistency. We validate the effectiveness of Q-Align through extensive experiments and analysis, and Q-Align outperforms state-of-the-art methods in appearance fidelity while maintaining competitive structure preservation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_21090 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment Kim, Namu Kweon, Wonbin Kim, Minsoo Yu, Hwanjo Computer Vision and Pattern Recognition We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by the Query-Key alignment. To tackle this issue, we introduce Q-Align, utilizing Query-Query alignment to mitigate attention leakage and improve the semantic alignment in zero-shot appearance transfer. Q-Align incorporates three core contributions: (1) Query-Query alignment, facilitating the sophisticated spatial semantic mapping between two images; (2) Key-Value rearrangement, enhancing feature correspondence through realignment; and (3) Attention refinement using rearranged keys and values to maintain semantic consistency. We validate the effectiveness of Q-Align through extensive experiments and analysis, and Q-Align outperforms state-of-the-art methods in appearance fidelity while maintaining competitive structure preservation. |
| title | Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.21090 |