Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Boxun, Du, Yuming, Liu, Zichang, Yang, Siyu, Jiang, Ziyang, Yan, Siqi, Saha, Rajasi, Pumarola, Albert, Wang, Wenchen, Li, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
by: Lv, Chengtao, et al.
Published: (2026)
by: Lv, Chengtao, et al.
Published: (2026)
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
by: Lapautre, Nicolas, et al.
Published: (2025)
by: Lapautre, Nicolas, et al.
Published: (2025)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
Autoregressive Distillation of Diffusion Transformers
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
by: Song, Enxin, et al.
Published: (2025)
by: Song, Enxin, et al.
Published: (2025)
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
by: Samuel, Dvir, et al.
Published: (2026)
by: Samuel, Dvir, et al.
Published: (2026)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
NOSA: Native and Offloadable Sparse Attention
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
by: Xiao, Steven, et al.
Published: (2025)
by: Xiao, Steven, et al.
Published: (2025)
Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models
by: Qiu, Junxiang, et al.
Published: (2026)
by: Qiu, Junxiang, et al.
Published: (2026)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
by: Liu, Kunhao, et al.
Published: (2025)
by: Liu, Kunhao, et al.
Published: (2025)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
by: Han, Shihao, et al.
Published: (2026)
by: Han, Shihao, et al.
Published: (2026)
SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks
by: Xu, Boxun, et al.
Published: (2025)
by: Xu, Boxun, et al.
Published: (2025)
DiffusionPoser: Real-time Human Motion Reconstruction From Arbitrary Sparse Sensors Using Autoregressive Diffusion
by: Van Wouwe, Tom, et al.
Published: (2023)
by: Van Wouwe, Tom, et al.
Published: (2023)
Bidirectional Sparse Attention for Faster Video Diffusion Training
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
by: Lu, Beijia, et al.
Published: (2025)
by: Lu, Beijia, et al.
Published: (2025)
Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
SparseD: Sparse Attention for Diffusion Language Models
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
VORTA: Efficient Video Diffusion via Routing Sparse Attention
by: Sun, Wenhao, et al.
Published: (2025)
by: Sun, Wenhao, et al.
Published: (2025)
LVSA: Training-Free Sparse Attention for Long Video Diffusion
by: Glorian, Gael, et al.
Published: (2026)
by: Glorian, Gael, et al.
Published: (2026)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
by: Xiu, Di, et al.
Published: (2025)
by: Xiu, Di, et al.
Published: (2025)
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
by: Long, Yongji, et al.
Published: (2026)
by: Long, Yongji, et al.
Published: (2026)
SneakPeek: Future-Guided Instructional Streaming Video Generation
by: Hong, Cheeun, et al.
Published: (2025)
by: Hong, Cheeun, et al.
Published: (2025)
Playing the Lottery With Concave Regularizers for Sparse Trainable Neural Networks
by: Fracastoro, Giulia, et al.
Published: (2025)
by: Fracastoro, Giulia, et al.
Published: (2025)
TabNSA: Native Sparse Attention for Efficient Tabular Data Learning
by: Eslamian, Ali, et al.
Published: (2025)
by: Eslamian, Ali, et al.
Published: (2025)
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
by: Xu, Boxun, et al.
Published: (2025)
by: Xu, Boxun, et al.
Published: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
by: Zhang, Hai, et al.
Published: (2026)
by: Zhang, Hai, et al.
Published: (2026)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Taming Teacher Forcing for Masked Autoregressive Video Generation
by: Zhou, Deyu, et al.
Published: (2025)
by: Zhou, Deyu, et al.
Published: (2025)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning
by: Ji, Kangye, et al.
Published: (2026)
by: Ji, Kangye, et al.
Published: (2026)
Similar Items
-
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025) -
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025) -
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
by: Lv, Chengtao, et al.
Published: (2026) -
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
by: Lapautre, Nicolas, et al.
Published: (2025) -
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)