Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yifan, Xiao, Zeqi, Wei, Tianyi, Yang, Shuai, Pan, Xingang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Video Diffusion Models are Training-free Motion Interpreter and Controller
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
Trajectory Attention for Fine-grained Video Motion Control
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
WorldMem: Long-term Consistent World Simulation with Memory
by: Xiao, Zeqi, et al.
Published: (2025)
by: Xiao, Zeqi, et al.
Published: (2025)
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
TokensGen: Harnessing Condensed Tokens for Long Video Generation
by: Ouyang, Wenqi, et al.
Published: (2025)
by: Ouyang, Wenqi, et al.
Published: (2025)
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
by: Wei, Tianyi, et al.
Published: (2025)
by: Wei, Tianyi, et al.
Published: (2025)
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
by: Wei, Tianyi, et al.
Published: (2024)
by: Wei, Tianyi, et al.
Published: (2024)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models
by: Fortes, Armando, et al.
Published: (2025)
by: Fortes, Armando, et al.
Published: (2025)
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
by: Li, Haopeng, et al.
Published: (2026)
by: Li, Haopeng, et al.
Published: (2026)
Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation
by: Xu, Boxun, et al.
Published: (2026)
by: Xu, Boxun, et al.
Published: (2026)
PI-Light: Physics-Inspired Diffusion for Full-Image Relighting
by: Liang, Zhexin, et al.
Published: (2026)
by: Liang, Zhexin, et al.
Published: (2026)
Boosting Monocular Metric Depth Estimation via Bokeh Rendering
by: Zhang, Hangwei, et al.
Published: (2025)
by: Zhang, Hangwei, et al.
Published: (2025)
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
by: Chen, Yongwei, et al.
Published: (2026)
by: Chen, Yongwei, et al.
Published: (2026)
SparseDiT: Token Sparsification for Efficient Diffusion Transformer
by: Chang, Shuning, et al.
Published: (2024)
by: Chang, Shuning, et al.
Published: (2024)
Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024)
by: Pu, Yifan, et al.
Published: (2024)
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
by: Bu, Jiazi, et al.
Published: (2026)
by: Bu, Jiazi, et al.
Published: (2026)
MvDrag3D: Drag-based Creative 3D Editing via Multi-view Generation-Reconstruction Priors
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
VORTA: Efficient Video Diffusion via Routing Sparse Attention
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
by: Shmilovich, Dor, et al.
Published: (2025)
by: Shmilovich, Dor, et al.
Published: (2025)
Filter-Guided Diffusion for Controllable Image Generation
by: Gu, Zeqi, et al.
Published: (2023)
by: Gu, Zeqi, et al.
Published: (2023)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
by: Wang, Yuxi, et al.
Published: (2026)
by: Wang, Yuxi, et al.
Published: (2026)
Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models
by: Xing, Zhening, et al.
Published: (2024)
by: Xing, Zhening, et al.
Published: (2024)
ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices
by: Du, Kunpeng, et al.
Published: (2026)
by: Du, Kunpeng, et al.
Published: (2026)
Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
by: Wang, Chung-Shien Brian, et al.
Published: (2025)
by: Wang, Chung-Shien Brian, et al.
Published: (2025)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
by: Xiao, Chaodong, et al.
Published: (2026)
by: Xiao, Chaodong, et al.
Published: (2026)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement
by: Luo, Yihang, et al.
Published: (2024)
by: Luo, Yihang, et al.
Published: (2024)
Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
by: Park, Dogyun, et al.
Published: (2025)
by: Park, Dogyun, et al.
Published: (2025)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
by: Ding, Hangliang, et al.
Published: (2025)
by: Ding, Hangliang, et al.
Published: (2025)
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
by: Ouyang, Wenqi, et al.
Published: (2024)
by: Ouyang, Wenqi, et al.
Published: (2024)
SP$^2$T: Sparse Proxy Attention for Dual-stream Point Transformer
by: Wan, Jiaxu, et al.
Published: (2024)
by: Wan, Jiaxu, et al.
Published: (2024)
Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention
by: Zhang, Wenhu, et al.
Published: (2026)
by: Zhang, Wenhu, et al.
Published: (2026)
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
by: Lan, Yushi, et al.
Published: (2025)
by: Lan, Yushi, et al.
Published: (2025)
Textured 3D Regenerative Morphing with 3D Diffusion Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Similar Items
-
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
by: Zhou, Yifan, et al.
Published: (2025) -
Video Diffusion Models are Training-free Motion Interpreter and Controller
by: Xiao, Zeqi, et al.
Published: (2024) -
Trajectory Attention for Fine-grained Video Motion Control
by: Xiao, Zeqi, et al.
Published: (2024) -
WorldMem: Long-term Consistent World Simulation with Memory
by: Xiao, Zeqi, et al.
Published: (2025) -
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)