Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gan, Chaofan, Zhao, Zicheng, Tu, Yuanpeng, Chen, Xi, Qin, Ziran, Chen, Tieyuan, Harandi, Mehrtash, Lin, Weiyao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915552700661760
author Gan, Chaofan
Zhao, Zicheng
Tu, Yuanpeng
Chen, Xi
Qin, Ziran
Chen, Tieyuan
Harandi, Mehrtash
Lin, Weiyao
author_facet Gan, Chaofan
Zhao, Zicheng
Tu, Yuanpeng
Chen, Xi
Qin, Ziran
Chen, Tieyuan
Harandi, Mehrtash
Lin, Weiyao
contents Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feature maps, yet their function remains poorly understood. In this work, we systematically investigate these activations to elucidate their role in visual generation. We found that these massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings. Importantly, our investigations further demonstrate that these massive activations play a key role in local detail synthesis, while having minimal impact on the overall semantic content of output. Building on these insights, we propose \textbf{D}etail \textbf{G}uidance (\textbf{DG}), a MAs-driven, training-free self-guidance strategy to explicitly enhance local detail fidelity for DiTs. Specifically, DG constructs a degraded ``detail-deficient'' model by disrupting MAs and leverages it to guide the original network toward higher-quality detail synthesis. Our DG can seamlessly integrate with Classifier-Free Guidance (CFG), enabling further refinements of fine-grained details. Extensive experiments demonstrate that our DG consistently improves fine-grained detail quality across various pre-trained DiTs (\eg, SD3, SD3.5, and Flux).
format Preprint
id arxiv_https___arxiv_org_abs_2510_11538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
Gan, Chaofan
Zhao, Zicheng
Tu, Yuanpeng
Chen, Xi
Qin, Ziran
Chen, Tieyuan
Harandi, Mehrtash
Lin, Weiyao
Computer Vision and Pattern Recognition
Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feature maps, yet their function remains poorly understood. In this work, we systematically investigate these activations to elucidate their role in visual generation. We found that these massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings. Importantly, our investigations further demonstrate that these massive activations play a key role in local detail synthesis, while having minimal impact on the overall semantic content of output. Building on these insights, we propose \textbf{D}etail \textbf{G}uidance (\textbf{DG}), a MAs-driven, training-free self-guidance strategy to explicitly enhance local detail fidelity for DiTs. Specifically, DG constructs a degraded ``detail-deficient'' model by disrupting MAs and leverages it to guide the original network toward higher-quality detail synthesis. Our DG can seamlessly integrate with Classifier-Free Guidance (CFG), enabling further refinements of fine-grained details. Extensive experiments demonstrate that our DG consistently improves fine-grained detail quality across various pre-trained DiTs (\eg, SD3, SD3.5, and Flux).
title Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.11538