Attention Sinks in Diffusion Transformers: A Causal Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Fangzheng, Summa, Brian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models
di: Wu, Fangzheng, et al.
Pubblicazione: (2026)
di: Wu, Fangzheng, et al.
Pubblicazione: (2026)
Model-Centric Diagnostics: A Framework for Internal State Readouts
di: Wu, Fangzheng, et al.
Pubblicazione: (2026)
di: Wu, Fangzheng, et al.
Pubblicazione: (2026)
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
di: Lu, Andrew, et al.
Pubblicazione: (2025)
di: Lu, Andrew, et al.
Pubblicazione: (2025)
Analysis of Attention in Video Diffusion Transformers
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
di: Liu, Xu, et al.
Pubblicazione: (2026)
di: Liu, Xu, et al.
Pubblicazione: (2026)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
di: Choi, Jiho, et al.
Pubblicazione: (2026)
di: Choi, Jiho, et al.
Pubblicazione: (2026)
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
di: Yoo, Suho, et al.
Pubblicazione: (2026)
di: Yoo, Suho, et al.
Pubblicazione: (2026)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
di: Feng, Wenfeng, et al.
Pubblicazione: (2025)
di: Feng, Wenfeng, et al.
Pubblicazione: (2025)
Causal Diffusion Transformers for Generative Modeling
di: Deng, Chaorui, et al.
Pubblicazione: (2024)
di: Deng, Chaorui, et al.
Pubblicazione: (2024)
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
di: Shmilovich, Dor, et al.
Pubblicazione: (2025)
di: Shmilovich, Dor, et al.
Pubblicazione: (2025)
Precipitation Nowcasting Using Diffusion Transformer with Causal Attention
di: Li, ChaoRong, et al.
Pubblicazione: (2024)
di: Li, ChaoRong, et al.
Pubblicazione: (2024)
Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning
di: Yi, Shuai, et al.
Pubblicazione: (2026)
di: Yi, Shuai, et al.
Pubblicazione: (2026)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
di: Xiao, Chaodong, et al.
Pubblicazione: (2026)
di: Xiao, Chaodong, et al.
Pubblicazione: (2026)
MSC: Multi-Scale Spatio-Temporal Causal Attention for Autoregressive Video Diffusion
di: Xu, Xunnong, et al.
Pubblicazione: (2024)
di: Xu, Xunnong, et al.
Pubblicazione: (2024)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
di: Ghafoorian, Mohsen, et al.
Pubblicazione: (2026)
di: Ghafoorian, Mohsen, et al.
Pubblicazione: (2026)
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
di: Li, Haopeng, et al.
Pubblicazione: (2026)
di: Li, Haopeng, et al.
Pubblicazione: (2026)
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
di: Zhou, Yifan, et al.
Pubblicazione: (2025)
di: Zhou, Yifan, et al.
Pubblicazione: (2025)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
di: Pu, Yifan, et al.
Pubblicazione: (2024)
di: Pu, Yifan, et al.
Pubblicazione: (2024)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
di: Yuan, Zhihang, et al.
Pubblicazione: (2024)
di: Yuan, Zhihang, et al.
Pubblicazione: (2024)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
di: Li, Haodong, et al.
Pubblicazione: (2026)
di: Li, Haodong, et al.
Pubblicazione: (2026)
An Analysis on Quantizing Diffusion Transformers
di: Yang, Yuewei, et al.
Pubblicazione: (2024)
di: Yang, Yuewei, et al.
Pubblicazione: (2024)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
di: Feng, Weilun, et al.
Pubblicazione: (2025)
di: Feng, Weilun, et al.
Pubblicazione: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
di: Liang, Cheng, et al.
Pubblicazione: (2026)
di: Liang, Cheng, et al.
Pubblicazione: (2026)
One Attention, One Scale: Phase-Aligned Rotary Positional Embeddings for Mixed-Resolution Diffusion Transformer
di: Wu, Haoyu, et al.
Pubblicazione: (2025)
di: Wu, Haoyu, et al.
Pubblicazione: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
di: Kang, Seil, et al.
Pubblicazione: (2025)
di: Kang, Seil, et al.
Pubblicazione: (2025)
SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
di: Rajabi, Javad, et al.
Pubblicazione: (2026)
di: Rajabi, Javad, et al.
Pubblicazione: (2026)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
di: Le, Minh Khoa, et al.
Pubblicazione: (2026)
di: Le, Minh Khoa, et al.
Pubblicazione: (2026)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
di: Ghafoorian, Mohsen, et al.
Pubblicazione: (2025)
di: Ghafoorian, Mohsen, et al.
Pubblicazione: (2025)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
di: Ding, Hangliang, et al.
Pubblicazione: (2025)
di: Ding, Hangliang, et al.
Pubblicazione: (2025)
Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer
di: Bui, Minh, et al.
Pubblicazione: (2024)
di: Bui, Minh, et al.
Pubblicazione: (2024)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
di: Fang, Tongcheng, et al.
Pubblicazione: (2026)
di: Fang, Tongcheng, et al.
Pubblicazione: (2026)
Vision Transformers with Hierarchical Attention
di: Liu, Yun, et al.
Pubblicazione: (2021)
di: Liu, Yun, et al.
Pubblicazione: (2021)
Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
di: Wang, Chung-Shien Brian, et al.
Pubblicazione: (2025)
di: Wang, Chung-Shien Brian, et al.
Pubblicazione: (2025)
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
di: Ye, Bo, et al.
Pubblicazione: (2026)
di: Ye, Bo, et al.
Pubblicazione: (2026)
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
di: Roy, Dip
Pubblicazione: (2025)
di: Roy, Dip
Pubblicazione: (2025)
Hadamard Attention Recurrent Transformer: A Strong Baseline for Stereo Matching Transformer
di: Chen, Ziyang, et al.
Pubblicazione: (2025)
di: Chen, Ziyang, et al.
Pubblicazione: (2025)
MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
di: Qiu, Qianru, et al.
Pubblicazione: (2025)
di: Qiu, Qianru, et al.
Pubblicazione: (2025)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
di: Wang, Zhongqi, et al.
Pubblicazione: (2025)
di: Wang, Zhongqi, et al.
Pubblicazione: (2025)
FG-TreeSeg: Flow-Guided Tree Crown Segmentation without Instance Annotations
di: Chen, Pengyu, et al.
Pubblicazione: (2026)
di: Chen, Pengyu, et al.
Pubblicazione: (2026)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
di: Luo, Jiayun, et al.
Pubblicazione: (2025)
di: Luo, Jiayun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models
di: Wu, Fangzheng, et al.
Pubblicazione: (2026) -
Model-Centric Diagnostics: A Framework for Internal State Readouts
di: Wu, Fangzheng, et al.
Pubblicazione: (2026) -
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
di: Lu, Andrew, et al.
Pubblicazione: (2025) -
Analysis of Attention in Video Diffusion Transformers
di: Wen, Yuxin, et al.
Pubblicazione: (2025) -
SinkTrack: Attention Sink based Context Anchoring for Large Language Models
di: Liu, Xu, et al.
Pubblicazione: (2026)