From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Yuxin, Collins, Maxwell D, Hu, Miao, Yang, Huanrui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FAR: Function-preserving Attention Replacement for IMC-friendly Inference
by: Ren, Yuxin, et al.
Published: (2025)
by: Ren, Yuxin, et al.
Published: (2025)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
by: Pang, Jing-Cheng, et al.
Published: (2021)
by: Pang, Jing-Cheng, et al.
Published: (2021)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
by: Xia, Xiaojie, et al.
Published: (2026)
by: Xia, Xiaojie, et al.
Published: (2026)
Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
by: Ashok, Arjun, et al.
Published: (2023)
by: Ashok, Arjun, et al.
Published: (2023)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
by: Zhang, Jinouwen, et al.
Published: (2024)
by: Zhang, Jinouwen, et al.
Published: (2024)
Online Policy Distillation with Decision-Attention
by: Yu, Xinqiang, et al.
Published: (2024)
by: Yu, Xinqiang, et al.
Published: (2024)
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow
by: Azarkhalili, Behrooz, et al.
Published: (2025)
by: Azarkhalili, Behrooz, et al.
Published: (2025)
Defending Membership Inference Attacks via Privacy-aware Sparsity Tuning
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
by: Draye, Florent, et al.
Published: (2025)
by: Draye, Florent, et al.
Published: (2025)
A Compression Perspective on Simplicity Bias
by: Marty, Tom, et al.
Published: (2026)
by: Marty, Tom, et al.
Published: (2026)
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
by: Gu, Youping, et al.
Published: (2025)
by: Gu, Youping, et al.
Published: (2025)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025)
by: Gao, Yizhao, et al.
Published: (2025)
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Empty SPACE: Cross-Attention Sparsity for Concept Erasure in Diffusion Models
by: Novello, Nicola, et al.
Published: (2026)
by: Novello, Nicola, et al.
Published: (2026)
Large Language Model Compression with Global Rank and Sparsity Optimization
by: Zhou, Changhai, et al.
Published: (2025)
by: Zhou, Changhai, et al.
Published: (2025)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
by: Molom-Ochir, Tergel, et al.
Published: (2024)
by: Molom-Ochir, Tergel, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
BONSAI: Bayesian Optimization with Natural Simplicity and Interpretability
by: Daulton, Samuel, et al.
Published: (2026)
by: Daulton, Samuel, et al.
Published: (2026)
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
by: Kim, Kwanyoung, et al.
Published: (2025)
by: Kim, Kwanyoung, et al.
Published: (2025)
Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
by: Mittal, Akshay, et al.
Published: (2025)
by: Mittal, Akshay, et al.
Published: (2025)
S2O: Early Stopping for Sparse Attention via Online Permutation
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Towards Robust Knowledge Tracing Models via k-Sparse Attention
by: Huang, Shuyan, et al.
Published: (2024)
by: Huang, Shuyan, et al.
Published: (2024)
PAT: Pruning-Aware Tuning for Large Language Models
by: Liu, Yijiang, et al.
Published: (2024)
by: Liu, Yijiang, et al.
Published: (2024)
IGANN Sparse: Bridging Sparsity and Interpretability with Non-linear Insight
by: Stoecker, Theodor, et al.
Published: (2024)
by: Stoecker, Theodor, et al.
Published: (2024)
Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation
by: Zhang, Xiaoxiong, et al.
Published: (2024)
by: Zhang, Xiaoxiong, et al.
Published: (2024)
Effective Interplay between Sparsity and Quantization: From Theory to Practice
by: Harma, Simla Burcu, et al.
Published: (2024)
by: Harma, Simla Burcu, et al.
Published: (2024)
Improving Sparse Autoencoder with Dynamic Attention
by: Wang, Dongsheng, et al.
Published: (2026)
by: Wang, Dongsheng, et al.
Published: (2026)
SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting
by: Ding, Hang, et al.
Published: (2025)
by: Ding, Hang, et al.
Published: (2025)
Similar Items
-
FAR: Function-preserving Attention Replacement for IMC-friendly Inference
by: Ren, Yuxin, et al.
Published: (2025) -
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
by: Pang, Jing-Cheng, et al.
Published: (2021) -
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024) -
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026) -
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
by: Xia, Xiaojie, et al.
Published: (2026)