Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gan, Chaofan, Tu, Yuanpeng, Chen, Xi, Chen, Tieyuan, Li, Yuxi, Harandi, Mehrtash, Lin, Weiyao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917068405735424
author Gan, Chaofan
Tu, Yuanpeng
Chen, Xi
Chen, Tieyuan
Li, Yuxi
Harandi, Mehrtash
Lin, Weiyao
author_facet Gan, Chaofan
Tu, Yuanpeng
Chen, Xi
Chen, Tieyuan
Li, Yuxi
Harandi, Mehrtash
Lin, Weiyao
contents Pre-trained stable diffusion models (SD) have shown great advances in visual correspondence. In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as \textit{massive activations}, leading to uninformative representations and significant performance degradation for DiTs. The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information. We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs. Building on these findings, we propose the \textbf{Di}ffusion \textbf{T}ransformer \textbf{F}eature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs. Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation. Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations. Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (\eg, with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.).
format Preprint
id arxiv_https___arxiv_org_abs_2505_18584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
Gan, Chaofan
Tu, Yuanpeng
Chen, Xi
Chen, Tieyuan
Li, Yuxi
Harandi, Mehrtash
Lin, Weiyao
Computer Vision and Pattern Recognition
Pre-trained stable diffusion models (SD) have shown great advances in visual correspondence. In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as \textit{massive activations}, leading to uninformative representations and significant performance degradation for DiTs. The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information. We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs. Building on these findings, we propose the \textbf{Di}ffusion \textbf{T}ransformer \textbf{F}eature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs. Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation. Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations. Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (\eg, with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.).
title Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18584