Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Zunhai, Zhang, Hengyuan, Wu, Wei, Zhang, Yifan, Liu, Yaxiu, Xiao, He, Yang, Qingyao, Sun, Yuxuan, Yang, Rui, Zhang, Chao, Fan, Keyu, Ye, Weihao, Xiong, Jing, Shen, Hui, Tao, Chaofan, Wu, Taiqiang, Wan, Zhongwei, Qian, Yulei, Xie, Yuchen, Wong, Ngai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024)
MMFormalizer: Multimodal Autoformalization in the Wild
von: Xiong, Jing, et al.
Veröffentlicht: (2026)
von: Xiong, Jing, et al.
Veröffentlicht: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
von: Xiong, Jing, et al.
Veröffentlicht: (2026)
von: Xiong, Jing, et al.
Veröffentlicht: (2026)
Revisiting Model Interpolation for Efficient Reasoning
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
The Art of Efficient Reasoning: Data, Reward, and Optimization
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
Mixture-of-Subspaces in Low-Rank Adaptation
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models
von: Chen, Junyu, et al.
Veröffentlicht: (2026)
von: Chen, Junyu, et al.
Veröffentlicht: (2026)
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)
Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
MINR: Efficient Implicit Neural Representations for Multi-Image Encoding
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
Timber: Training-free Instruct Model Refining with Base via Effective Rank
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
DoPE: Denoising Rotary Position Embedding
von: Xiong, Jing, et al.
Veröffentlicht: (2025)
von: Xiong, Jing, et al.
Veröffentlicht: (2025)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
von: Xiao, He, et al.
Veröffentlicht: (2025)
von: Xiao, He, et al.
Veröffentlicht: (2025)
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2023)
CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
On the Existence and Behavior of Secondary Attention Sinks
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural Representations
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
von: Chen, Xinrong, et al.
Veröffentlicht: (2026)
von: Chen, Xinrong, et al.
Veröffentlicht: (2026)
Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
Autoregressive Models in Vision: A Survey
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
von: Shen, Hui, et al.
Veröffentlicht: (2025)
von: Shen, Hui, et al.
Veröffentlicht: (2025)
LoCa: Logit Calibration for Knowledge Distillation
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
von: Yang, Zebin, et al.
Veröffentlicht: (2024)
von: Yang, Zebin, et al.
Veröffentlicht: (2024)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
Attention Sinks and Outliers in Attention Residuals
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
Re-Activating Frozen Primitives for 3D Gaussian Splatting
von: Cheng, Yuxin, et al.
Veröffentlicht: (2025)
von: Cheng, Yuxin, et al.
Veröffentlicht: (2025)
QuadINR: Hardware-Efficient Implicit Neural Representations Through Quadratic Activation
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024) -
MMFormalizer: Multimodal Autoformalization in the Wild
von: Xiong, Jing, et al.
Veröffentlicht: (2026) -
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026) -
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026) -
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025)