Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Stephen, Khan, Mustafa, Papyan, Vardan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
por: Zhang, Stephen, et al.
Publicado: (2024)
por: Zhang, Stephen, et al.
Publicado: (2024)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
por: Zhang, Stephen, et al.
Publicado: (2024)
por: Zhang, Stephen, et al.
Publicado: (2024)
Transformer Block Coupling and its Correlation with Generalization in LLMs
por: Aubry, Murdock, et al.
Publicado: (2024)
por: Aubry, Murdock, et al.
Publicado: (2024)
On the Existence and Behavior of Secondary Attention Sinks
por: Wong, Jeffrey T. H., et al.
Publicado: (2025)
por: Wong, Jeffrey T. H., et al.
Publicado: (2025)
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)
por: Gu, Xiangming, et al.
Publicado: (2024)
Sink-Aware Pruning for Diffusion Language Models
por: Myrzakhan, Aidar, et al.
Publicado: (2026)
por: Myrzakhan, Aidar, et al.
Publicado: (2026)
Linguistic Collapse: Neural Collapse in (Large) Language Models
por: Wu, Robert, et al.
Publicado: (2024)
por: Wu, Robert, et al.
Publicado: (2024)
Limitations of Normalization in Attention Mechanism
por: Mudarisov, Timur, et al.
Publicado: (2025)
por: Mudarisov, Timur, et al.
Publicado: (2025)
Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
por: Mamtani, Sumit, et al.
Publicado: (2025)
por: Mamtani, Sumit, et al.
Publicado: (2025)
Residual Alignment: Uncovering the Mechanisms of Residual Networks
por: Li, Jianing, et al.
Publicado: (2024)
por: Li, Jianing, et al.
Publicado: (2024)
A Simple Attention-Based Mechanism for Bimodal Emotion Classification
por: Elabd, Mazen, et al.
Publicado: (2024)
por: Elabd, Mazen, et al.
Publicado: (2024)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
por: Shin, Seungjun, et al.
Publicado: (2025)
por: Shin, Seungjun, et al.
Publicado: (2025)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
por: Zhu, Shiyi, et al.
Publicado: (2023)
por: Zhu, Shiyi, et al.
Publicado: (2023)
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
por: Jiang, Yuhang
Publicado: (2026)
por: Jiang, Yuhang
Publicado: (2026)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
por: Kim, Jongwook, et al.
Publicado: (2025)
por: Kim, Jongwook, et al.
Publicado: (2025)
Tag-LLM: Repurposing General-Purpose LLMs for Specialized Domains
por: Shen, Junhong, et al.
Publicado: (2024)
por: Shen, Junhong, et al.
Publicado: (2024)
Ensemble Learning for Large Language Models in Text and Code Generation: A Survey
por: Ashiga, Mari, et al.
Publicado: (2025)
por: Ashiga, Mari, et al.
Publicado: (2025)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
por: Riachi, Roland, et al.
Publicado: (2025)
por: Riachi, Roland, et al.
Publicado: (2025)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
por: Mirtaheri, Parsa, et al.
Publicado: (2026)
por: Mirtaheri, Parsa, et al.
Publicado: (2026)
JTCSE: Joint Tensor-Modulus Constraints and Cross-Attention for Unsupervised Contrastive Learning of Sentence Embeddings
por: Zong, Tianyu, et al.
Publicado: (2025)
por: Zong, Tianyu, et al.
Publicado: (2025)
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
por: Leemann, Tobias, et al.
Publicado: (2024)
por: Leemann, Tobias, et al.
Publicado: (2024)
Detecting Suicidal Ideation in Text with Interpretable Deep Learning: A CNN-BiGRU with Attention Mechanism
por: Bhuiyan, Mohaiminul Islam, et al.
Publicado: (2025)
por: Bhuiyan, Mohaiminul Islam, et al.
Publicado: (2025)
SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
por: Pradhan, Bidyapati, et al.
Publicado: (2025)
por: Pradhan, Bidyapati, et al.
Publicado: (2025)
Attention Sinks and Outliers in Attention Residuals
por: Luo, Haozheng, et al.
Publicado: (2026)
por: Luo, Haozheng, et al.
Publicado: (2026)
TopicTag: Automatic Annotation of NMF Topic Models Using Chain of Thought and Prompt Tuning with LLMs
por: Wanna, Selma, et al.
Publicado: (2024)
por: Wanna, Selma, et al.
Publicado: (2024)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
por: Fu, Zichuan, et al.
Publicado: (2026)
por: Fu, Zichuan, et al.
Publicado: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
por: Yuan, Jingyang, et al.
Publicado: (2025)
por: Yuan, Jingyang, et al.
Publicado: (2025)
Higher-order Linear Attention
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Laplacian Heads Improve Transformers by Smoothing Token Representations
por: Zhang, Yuchong, et al.
Publicado: (2026)
por: Zhang, Yuchong, et al.
Publicado: (2026)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
por: Sun, Shangwen, et al.
Publicado: (2026)
por: Sun, Shangwen, et al.
Publicado: (2026)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
por: Wang, Yiming, et al.
Publicado: (2024)
por: Wang, Yiming, et al.
Publicado: (2024)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
por: Zhang, Tianyi, et al.
Publicado: (2024)
por: Zhang, Tianyi, et al.
Publicado: (2024)
Mechanics of Next Token Prediction with Self-Attention
por: Li, Yingcong, et al.
Publicado: (2024)
por: Li, Yingcong, et al.
Publicado: (2024)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
por: Xie, Linxi, et al.
Publicado: (2025)
por: Xie, Linxi, et al.
Publicado: (2025)
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
por: Chen, Jianlv, et al.
Publicado: (2024)
por: Chen, Jianlv, et al.
Publicado: (2024)
From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
por: Khan, Imran
Publicado: (2025)
por: Khan, Imran
Publicado: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
CL4KGE: A Curriculum Learning Method for Knowledge Graph Embedding
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
Ejemplares similares
-
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
por: Zhang, Stephen, et al.
Publicado: (2024) -
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
por: Zhang, Stephen, et al.
Publicado: (2024) -
Transformer Block Coupling and its Correlation with Generalization in LLMs
por: Aubry, Murdock, et al.
Publicado: (2024) -
On the Existence and Behavior of Secondary Attention Sinks
por: Wong, Jeffrey T. H., et al.
Publicado: (2025) -
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)