Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ko, Jungmin, Park, Jungwon, Kim, Jimyeong, Choi, Changin, Lee, Wonseok, Rhee, Wonjong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911726641872896
author Ko, Jungmin
Park, Jungwon
Kim, Jimyeong
Choi, Changin
Lee, Wonseok
Rhee, Wonjong
author_facet Ko, Jungmin
Park, Jungwon
Kim, Jimyeong
Choi, Changin
Lee, Wonseok
Rhee, Wonjong
contents Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fundamentally challenging problem. Existing approaches, including prompt negation, post-hoc editing, and negative guidance, remain insufficient for explicit concept suppression, often failing to remove the target concept or degrading overall image quality. To this end, we propose Orthogonal Negative Guidance in attention feature space, a training-free method that operates in the attention output space of MM-DiT-based T2I transformers. Our method orthogonalizes negative-prompt attention features with respect to positive-prompt features and subtracts only the orthogonal component, suppressing unwanted concepts while preserving desired semantics. Experiments on FLUX-dev and FLUX-schnell show that our method achieves favorable trade-offs between concept suppression, prompt alignment, and image quality. In human evaluation, our method outperforms the second-best baseline by 18.78%. We further show that our method supports multi-concept suppression and adjustable concept suppression.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29390
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
Ko, Jungmin
Park, Jungwon
Kim, Jimyeong
Choi, Changin
Lee, Wonseok
Rhee, Wonjong
Computer Vision and Pattern Recognition
Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fundamentally challenging problem. Existing approaches, including prompt negation, post-hoc editing, and negative guidance, remain insufficient for explicit concept suppression, often failing to remove the target concept or degrading overall image quality. To this end, we propose Orthogonal Negative Guidance in attention feature space, a training-free method that operates in the attention output space of MM-DiT-based T2I transformers. Our method orthogonalizes negative-prompt attention features with respect to positive-prompt features and subtracts only the orthogonal component, suppressing unwanted concepts while preserving desired semantics. Experiments on FLUX-dev and FLUX-schnell show that our method achieves favorable trade-offs between concept suppression, prompt alignment, and image quality. In human evaluation, our method outperforms the second-best baseline by 18.78%. We further show that our method supports multi-concept suppression and adjustable concept suppression.
title Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.29390