SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jintao, Wang, Haoxu, Jiang, Kai, Yang, Shuo, Zheng, Kaiwen, Xi, Haocheng, Wang, Ziteng, Zhu, Hongzhou, Zhao, Min, Stoica, Ion, Gonzalez, Joseph E., Zhu, Jun, Chen, Jianfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SageBwd: A Trainable Low-bit Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
von: Zhou, Xuanyi, et al.
Veröffentlicht: (2026)
von: Zhou, Xuanyi, et al.
Veröffentlicht: (2026)
HashAttention: Semantic Sparsity for Faster Inference
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
vAttention: Verified Sparse Attention
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
Oscillation-Reduced MXFP4 Training for Vision Transformers
von: Chen, Yuxiang, et al.
Veröffentlicht: (2025)
von: Chen, Yuxiang, et al.
Veröffentlicht: (2025)
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
von: Wang, Ziteng, et al.
Veröffentlicht: (2024)
von: Wang, Ziteng, et al.
Veröffentlicht: (2024)
Efficient Backpropagation with Variance-Controlled Adaptive Sampling
von: Wang, Ziteng, et al.
Veröffentlicht: (2024)
von: Wang, Ziteng, et al.
Veröffentlicht: (2024)
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
VSA: Faster Video Diffusion with Trainable Sparse Attention
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2025)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block Quantization
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025)
von: Zhao, Min, et al.
Veröffentlicht: (2025)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025)
von: Zhao, Min, et al.
Veröffentlicht: (2025)
Some Present-Day Problems of Romanian Library Science
von: Stoica, Ion
Veröffentlicht: (1973)
von: Stoica, Ion
Veröffentlicht: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
von: Stoica, Ion
Veröffentlicht: (1972)
von: Stoica, Ion
Veröffentlicht: (1972)
SparseDM: Toward Sparse Efficient Diffusion Models
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
von: Ding, Hangliang, et al.
Veröffentlicht: (2025)
von: Ding, Hangliang, et al.
Veröffentlicht: (2025)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
Accelerating Transformer Pre-training with 2:4 Sparsity
von: Hu, Yuezhou, et al.
Veröffentlicht: (2024)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
von: Zhao, Min, et al.
Veröffentlicht: (2024)
von: Zhao, Min, et al.
Veröffentlicht: (2024)
Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2023)
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2023)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Inference Time Context Sparsity: Illusion or Opportunity?
von: Joshi, Sahil, et al.
Veröffentlicht: (2026)
von: Joshi, Sahil, et al.
Veröffentlicht: (2026)
Scaling Linear Attention with Sparse State Expansion
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
Consistency Diffusion Bridge Models
von: He, Guande, et al.
Veröffentlicht: (2024)
von: He, Guande, et al.
Veröffentlicht: (2024)
Diffusion Bridge Implicit Models
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zheng, Kaiwen, et al.
Veröffentlicht: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026) -
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
SageBwd: A Trainable Low-bit Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2026) -
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025) -
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)