StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jeong, Jaeseok, Kim, Junho, Lee, Gayoung, Choi, Yunjey, Uh, Youngjung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909831091191808
author Jeong, Jaeseok
Kim, Junho
Lee, Gayoung
Choi, Yunjey
Uh, Youngjung
author_facet Jeong, Jaeseok
Kim, Junho
Lee, Gayoung
Choi, Yunjey
Uh, Youngjung
contents In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and content. However, existing methods often suffer from content leakage, where undesired elements of the visual style prompt are transferred along with the intended style. To address this issue, we 1) extend classifier-free guidance (CFG) to utilize swapping self-attention and propose 2) negative visual query guidance (NVQG) to reduce the transfer of unwanted contents. NVQG employs negative score by intentionally simulating content leakage scenarios that swap queries instead of key and values of self-attention layers from visual style prompts. This simple yet effective method significantly reduces content leakage. Furthermore, we provide careful solutions for using a real image as visual style prompts. Through extensive evaluation across various styles and text prompts, our method demonstrates superiority over existing approaches, reflecting the style of the references, and ensuring that resulting images match the text prompts. Our code is available \href{https://github.com/naver-ai/StyleKeeper}{here}.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
Jeong, Jaeseok
Kim, Junho
Lee, Gayoung
Choi, Yunjey
Uh, Youngjung
Computer Vision and Pattern Recognition
In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and content. However, existing methods often suffer from content leakage, where undesired elements of the visual style prompt are transferred along with the intended style. To address this issue, we 1) extend classifier-free guidance (CFG) to utilize swapping self-attention and propose 2) negative visual query guidance (NVQG) to reduce the transfer of unwanted contents. NVQG employs negative score by intentionally simulating content leakage scenarios that swap queries instead of key and values of self-attention layers from visual style prompts. This simple yet effective method significantly reduces content leakage. Furthermore, we provide careful solutions for using a real image as visual style prompts. Through extensive evaluation across various styles and text prompts, our method demonstrates superiority over existing approaches, reflecting the style of the references, and ensuring that resulting images match the text prompts. Our code is available \href{https://github.com/naver-ai/StyleKeeper}{here}.
title StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.06827