Taming Outlier Tokens in Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiaoyu, Wang, Yifei, Fu, Tsu-Jui, Chen, Liang-Chieh, Gan, Zhe, Wei, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Latent Diffusion Model for Neural Radiance Field Inpainting
by: Lin, Chieh Hubert, et al.
Published: (2024)
by: Lin, Chieh Hubert, et al.
Published: (2024)
On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Taming Diffusion for Dataset Distillation with High Representativeness
by: Zhao, Lin, et al.
Published: (2025)
by: Zhao, Lin, et al.
Published: (2025)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
Accelerating Diffusion Transformers with Token-wise Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Neural Isometries: Taming Transformations for Equivariant ML
by: Mitchel, Thomas W., et al.
Published: (2024)
by: Mitchel, Thomas W., et al.
Published: (2024)
Taming Score-Based Denoisers in ADMM: A Convergent Plug-and-Play Framework
by: Shrestha, Rajesh, et al.
Published: (2026)
by: Shrestha, Rajesh, et al.
Published: (2026)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024)
by: Wu, Xiaoping, et al.
Published: (2024)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
FSL-Rectifier: Rectify Outliers in Few-Shot Learning via Test-Time Augmentation
by: Bai, Yunwei, et al.
Published: (2024)
by: Bai, Yunwei, et al.
Published: (2024)
MeanFlow Transformers with Representation Autoencoders
by: Hu, Zheyuan, et al.
Published: (2025)
by: Hu, Zheyuan, et al.
Published: (2025)
Leveraging Model Guidance to Extract Training Data from Personalized Diffusion Models
by: Wu, Xiaoyu, et al.
Published: (2024)
by: Wu, Xiaoyu, et al.
Published: (2024)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
by: Hong, Susung, et al.
Published: (2025)
by: Hong, Susung, et al.
Published: (2025)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
High-Order Matching for One-Step Shortcut Diffusion Models
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
FrameBridge: Improving Image-to-Video Generation with Bridge Models
by: Wang, Yuji, et al.
Published: (2024)
by: Wang, Yuji, et al.
Published: (2024)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
Robustness Tokens: Towards Adversarial Robustness of Transformers
by: Pulfer, Brian, et al.
Published: (2025)
by: Pulfer, Brian, et al.
Published: (2025)
Scaling Diffusion Transformers Efficiently via $μ$P
by: Zheng, Chenyu, et al.
Published: (2025)
by: Zheng, Chenyu, et al.
Published: (2025)
OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
by: Chu, Huanpeng, et al.
Published: (2025)
by: Chu, Huanpeng, et al.
Published: (2025)
Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning
by: Jeong, Wooseong, et al.
Published: (2025)
by: Jeong, Wooseong, et al.
Published: (2025)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
by: Yang, Yaoxin, et al.
Published: (2025)
by: Yang, Yaoxin, et al.
Published: (2025)
H$_{2}$OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation
by: Li, Wenhao, et al.
Published: (2023)
by: Li, Wenhao, et al.
Published: (2023)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
by: Chiang, Jui-Che, et al.
Published: (2024)
by: Chiang, Jui-Che, et al.
Published: (2024)
Non-Linear Outlier Synthesis for Out-of-Distribution Detection
by: Doorenbos, Lars, et al.
Published: (2024)
by: Doorenbos, Lars, et al.
Published: (2024)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
CAPM: Fast and Robust Verification on Maxpool-based CNN via Dual Network
by: Bai, Jia-Hau, et al.
Published: (2024)
by: Bai, Jia-Hau, et al.
Published: (2024)
Feature Purified Transformer With Cross-level Feature Guiding Decoder For Multi-class OOD and Anomaly Deteciton
by: Lin, Jerry Chun-Wei, et al.
Published: (2024)
by: Lin, Jerry Chun-Wei, et al.
Published: (2024)
TinyFusion: Diffusion Transformers Learned Shallow
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Similar Items
-
Taming Latent Diffusion Model for Neural Radiance Field Inpainting
by: Lin, Chieh Hubert, et al.
Published: (2024) -
On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024) -
Taming Diffusion for Dataset Distillation with High Representativeness
by: Zhao, Lin, et al.
Published: (2025) -
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
by: You, Haoran, et al.
Published: (2024) -
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)