Attention Condensation via Sparsity Induced Regularized Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sason, Eli, Frolova, Darya, Nazarov, Boris, Goldberd, Felix |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Data: Towards Better Performing Domain-Specific Small Language Models
von: Nazarov, Boris, et al.
Veröffentlicht: (2025)
von: Nazarov, Boris, et al.
Veröffentlicht: (2025)
Crisp Attention: Regularizing Transformers via Structured Sparsity
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025)
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
von: Deng, Keqi, et al.
Veröffentlicht: (2026)
von: Deng, Keqi, et al.
Veröffentlicht: (2026)
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models
von: Apartsin, Alexander, et al.
Veröffentlicht: (2026)
von: Apartsin, Alexander, et al.
Veröffentlicht: (2026)
Sparsity-Accelerated Training for Large Language Models
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
von: Luo, Jiayun, et al.
Veröffentlicht: (2024)
von: Luo, Jiayun, et al.
Veröffentlicht: (2024)
RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts
von: Kharlamova, Darya, et al.
Veröffentlicht: (2026)
von: Kharlamova, Darya, et al.
Veröffentlicht: (2026)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
Training-Free Activation Sparsity in Large Language Models
von: Liu, James, et al.
Veröffentlicht: (2024)
von: Liu, James, et al.
Veröffentlicht: (2024)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
Accelerating Prefilling via Decoding-time Contribution Sparsity
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
On Strongly Regular Graphs and the Friendship Theorem
von: Sason, Igal
Veröffentlicht: (2025)
von: Sason, Igal
Veröffentlicht: (2025)
Sparsity Induction for Accurate Post-Training Pruning of Large Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2026)
von: Jiang, Minhao, et al.
Veröffentlicht: (2026)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
von: Huang, Weiyu, et al.
Veröffentlicht: (2025)
von: Huang, Weiyu, et al.
Veröffentlicht: (2025)
Regularization, Semi-supervision, and Supervision for a Plausible Attention-Based Explanation
von: Nguyen, Duc Hau, et al.
Veröffentlicht: (2025)
von: Nguyen, Duc Hau, et al.
Veröffentlicht: (2025)
Robust Training of Neural Networks at Arbitrary Precision and Sparsity
von: Ye, Chengxi, et al.
Veröffentlicht: (2024)
von: Ye, Chengxi, et al.
Veröffentlicht: (2024)
LongAttn: Selecting Long-context Training Data via Token-level Attention
von: Wu, Longyun, et al.
Veröffentlicht: (2025)
von: Wu, Longyun, et al.
Veröffentlicht: (2025)
LEWIS (LayEr WIse Sparsity) -- A Training Free Guided Model Merging Approach
von: Chopra, Hetarth, et al.
Veröffentlicht: (2025)
von: Chopra, Hetarth, et al.
Veröffentlicht: (2025)
Empowering Interdisciplinary Research with BERT-Based Models: An Approach Through SciBERT-CNN with Topic Modeling
von: Likhareva, Darya, et al.
Veröffentlicht: (2024)
von: Likhareva, Darya, et al.
Veröffentlicht: (2024)
Long Context Pre-Training with Lighthouse Attention
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
von: Chen, Yao, et al.
Veröffentlicht: (2026)
von: Chen, Yao, et al.
Veröffentlicht: (2026)
Sirius: Contextual Sparsity with Correction for Efficient LLMs
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis
von: Song, Zichen, et al.
Veröffentlicht: (2024)
von: Song, Zichen, et al.
Veröffentlicht: (2024)
Controllable Stylistic Text Generation with Train-Time Attribute-Regularized Diffusion
von: Zhou, Fan, et al.
Veröffentlicht: (2025)
von: Zhou, Fan, et al.
Veröffentlicht: (2025)
Lag-Relative Sparse Attention In Long Context Training
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
When Does Sparsity Mitigate the Curse of Depth in LLMs
von: Muhtar, Dilxat, et al.
Veröffentlicht: (2026)
von: Muhtar, Dilxat, et al.
Veröffentlicht: (2026)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
Attention Instruction: Amplifying Attention in the Middle via Prompting
von: Zhang, Meiru, et al.
Veröffentlicht: (2024)
von: Zhang, Meiru, et al.
Veröffentlicht: (2024)
Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity
von: Song, Zichen, et al.
Veröffentlicht: (2024)
von: Song, Zichen, et al.
Veröffentlicht: (2024)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
von: Liu, Xiaoran, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2024)
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
von: Xu, Chi, et al.
Veröffentlicht: (2025)
von: Xu, Chi, et al.
Veröffentlicht: (2025)
BeLink: Biomedical Entity Linking Meets Generative Re-Ranking
von: Shlyk, Darya, et al.
Veröffentlicht: (2026)
von: Shlyk, Darya, et al.
Veröffentlicht: (2026)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
von: You, Zeng, et al.
Veröffentlicht: (2025)
von: You, Zeng, et al.
Veröffentlicht: (2025)
Symmetric Dot-Product Attention for Efficient Training of BERT Language Models
von: Courtois, Martin, et al.
Veröffentlicht: (2024)
von: Courtois, Martin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rethinking Data: Towards Better Performing Domain-Specific Small Language Models
von: Nazarov, Boris, et al.
Veröffentlicht: (2025) -
Crisp Attention: Regularizing Transformers via Structured Sparsity
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025) -
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024) -
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
von: Deng, Keqi, et al.
Veröffentlicht: (2026) -
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)