Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahn, Donghoon, Kang, Jiwon, Lee, Sanghyun, Kim, Minjae, Min, Jaewon, Jang, Wooseok, Lee, Sangwu, Paul, Sayak, Hong, Susung, Kim, Seungryong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918181149343744
author Ahn, Donghoon
Kang, Jiwon
Lee, Sanghyun
Kim, Minjae
Min, Jaewon
Jang, Wooseok
Lee, Sangwu
Paul, Sayak
Hong, Susung
Kim, Seungryong
author_facet Ahn, Donghoon
Kang, Jiwon
Lee, Sanghyun
Kim, Minjae
Min, Jaewon
Jang, Wooseok
Lee, Sangwu
Paul, Sayak
Hong, Susung
Kim, Seungryong
contents Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose "HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head's attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection. We validate our method on modern large-scale DiT-based text-to-image models including Stable Diffusion 3 and FLUX.1, demonstrating superior performance in both general quality enhancement and style-specific guidance. Our work provides the first head-level analysis of attention perturbation in diffusion models, uncovering interpretable specialization within attention layers and enabling practical design of effective perturbation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10978
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models
Ahn, Donghoon
Kang, Jiwon
Lee, Sanghyun
Kim, Minjae
Min, Jaewon
Jang, Wooseok
Lee, Sangwu
Paul, Sayak
Hong, Susung
Kim, Seungryong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose "HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head's attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection. We validate our method on modern large-scale DiT-based text-to-image models including Stable Diffusion 3 and FLUX.1, demonstrating superior performance in both general quality enhancement and style-specific guidance. Our work provides the first head-level analysis of attention perturbation in diffusion models, uncovering interpretable specialization within attention layers and enabling practical design of effective perturbation strategies.
title Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.10978