Taming Outlier Tokens in Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiaoyu, Wang, Yifei, Fu, Tsu-Jui, Chen, Liang-Chieh, Gan, Zhe, Wei, Chen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913095684718592
author Wu, Xiaoyu
Wang, Yifei
Fu, Tsu-Jui
Chen, Liang-Chieh
Gan, Zhe
Wei, Chen
author_facet Wu, Xiaoyu
Wang, Yifei
Fu, Tsu-Jui
Chen, Liang-Chieh
Gan, Zhe
Wei, Chen
contents We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models remains underexplored. We show that this phenomenon appears in both the encoder and denoiser of modern Representation Autoencoder (RAE)-DiT pipelines: pretrained ViT encoders can produce outlier representations, and DiTs themselves can develop internal outlier tokens, especially in intermediate layers. Moreover, simply masking high-norm tokens does not improve performance, indicating that the problem is not only caused by a few extreme values, but is more closely related to corrupted local patch semantics. To address this issue, we introduce Dual-Stage Registers (DSR), a register-based intervention for both components: trained registers when available, recursive test-time registers otherwise, and diffusion registers for the denoiser. Across ImageNet and large-scale text-to-image generation, these interventions consistently reduce outlier artifacts and improve generation quality. Our results highlight outlier-token control as an important ingredient in building stronger DiTs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05206
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Taming Outlier Tokens in Diffusion Transformers
Wu, Xiaoyu
Wang, Yifei
Fu, Tsu-Jui
Chen, Liang-Chieh
Gan, Zhe
Wei, Chen
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models remains underexplored. We show that this phenomenon appears in both the encoder and denoiser of modern Representation Autoencoder (RAE)-DiT pipelines: pretrained ViT encoders can produce outlier representations, and DiTs themselves can develop internal outlier tokens, especially in intermediate layers. Moreover, simply masking high-norm tokens does not improve performance, indicating that the problem is not only caused by a few extreme values, but is more closely related to corrupted local patch semantics. To address this issue, we introduce Dual-Stage Registers (DSR), a register-based intervention for both components: trained registers when available, recursive test-time registers otherwise, and diffusion registers for the denoiser. Across ImageNet and large-scale text-to-image generation, these interventions consistently reduce outlier artifacts and improve generation quality. Our results highlight outlier-token control as an important ingredient in building stronger DiTs.
title Taming Outlier Tokens in Diffusion Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.05206