Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Jiayi, Chen, Jiayu, Wang, Jiankun, Wang, Cong, Zhu, Hanxin, Sun, Qingyun, Gao, Chen, Chen, Zhibo, Li, Jianxin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
by: Luo, Jiayi, et al.
Published: (2026)
by: Luo, Jiayi, et al.
Published: (2026)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
Sparse Recovery for Holographic MIMO Channels: Leveraging the Clustered Sparsity
by: Guo, Yuqing, et al.
Published: (2024)
by: Guo, Yuqing, et al.
Published: (2024)
STS: Efficient Sparse Attention with Speculative Token Sparsity
by: Xu, Ceyu, et al.
Published: (2026)
by: Xu, Ceyu, et al.
Published: (2026)
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
by: Liu, Xuewen, et al.
Published: (2025)
by: Liu, Xuewen, et al.
Published: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
by: Zhu, Hanxin, et al.
Published: (2026)
by: Zhu, Hanxin, et al.
Published: (2026)
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
by: Deng, Keqi, et al.
Published: (2026)
by: Deng, Keqi, et al.
Published: (2026)
Attention Condensation via Sparsity Induced Regularized Training
by: Sason, Eli, et al.
Published: (2025)
by: Sason, Eli, et al.
Published: (2025)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
by: Lu, Beijia, et al.
Published: (2025)
by: Lu, Beijia, et al.
Published: (2025)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Spark Transformer: Reactivating Sparsity in FFN and Attention
by: You, Chong, et al.
Published: (2025)
by: You, Chong, et al.
Published: (2025)
Dynamic Sparse Training with Structured Sparsity
by: Lasby, Mike, et al.
Published: (2023)
by: Lasby, Mike, et al.
Published: (2023)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
by: Fang, Tongcheng, et al.
Published: (2026)
by: Fang, Tongcheng, et al.
Published: (2026)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
by: Xiao, Chaodong, et al.
Published: (2026)
by: Xiao, Chaodong, et al.
Published: (2026)
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
by: Yuan, Jiayi, et al.
Published: (2025)
by: Yuan, Jiayi, et al.
Published: (2025)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025)
by: Liu, Akide, et al.
Published: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Light Field Compression Based on Implicit Neural Representation
by: Wang, Henan, et al.
Published: (2024)
by: Wang, Henan, et al.
Published: (2024)
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
by: Ren, Yuxin, et al.
Published: (2026)
by: Ren, Yuxin, et al.
Published: (2026)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
by: Lin, Chaofan, et al.
Published: (2025)
by: Lin, Chaofan, et al.
Published: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
by: Fan, Zhenkun, et al.
Published: (2026)
by: Fan, Zhenkun, et al.
Published: (2026)
Crisp Attention: Regularizing Transformers via Structured Sparsity
by: Gandhi, Sagar, et al.
Published: (2025)
by: Gandhi, Sagar, et al.
Published: (2025)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
by: Tan, Xin, et al.
Published: (2025)
by: Tan, Xin, et al.
Published: (2025)
Discrepancy Minimization in Input-Sparsity Time
by: Deng, Yichuan, et al.
Published: (2022)
by: Deng, Yichuan, et al.
Published: (2022)
Sparsity-Accelerated Training for Large Language Models
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
LVSA: Training-Free Sparse Attention for Long Video Diffusion
by: Glorian, Gael, et al.
Published: (2026)
by: Glorian, Gael, et al.
Published: (2026)
Embody4D: A Generalist 4D World Model for Embodied AI
by: Tu, Peiyan, et al.
Published: (2026)
by: Tu, Peiyan, et al.
Published: (2026)
Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity
by: Fasina, Oluwadamilola, et al.
Published: (2025)
by: Fasina, Oluwadamilola, et al.
Published: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
CASP: Compression of Large Multimodal Models Based on Attention Sparsity
by: Gholami, Mohsen, et al.
Published: (2025)
by: Gholami, Mohsen, et al.
Published: (2025)
Similar Items
-
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
by: Luo, Jiayi, et al.
Published: (2026) -
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024) -
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026) -
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026) -
Sparse Recovery for Holographic MIMO Channels: Leveraging the Clustered Sparsity
by: Guo, Yuqing, et al.
Published: (2024)