Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Xingyue, Ding, Xueying, Ju, Mingxuan, Liu, Yozen, Shah, Neil, Zhao, Tong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
por: Ding, Xueying, et al.
Publicado: (2025)
por: Ding, Xueying, et al.
Publicado: (2025)
How Does Message Passing Improve Collaborative Filtering?
por: Ju, Mingxuan, et al.
Publicado: (2024)
por: Ju, Mingxuan, et al.
Publicado: (2024)
Robust Training Objectives Improve Embedding-based Retrieval in Industrial Recommendation Systems
por: Kolodner, Matthew, et al.
Publicado: (2024)
por: Kolodner, Matthew, et al.
Publicado: (2024)
Node Duplication Improves Cold-start Link Prediction
por: Guo, Zhichun, et al.
Publicado: (2024)
por: Guo, Zhichun, et al.
Publicado: (2024)
On the Role of Weight Decay in Collaborative Filtering: A Popularity Perspective
por: Loveland, Donald, et al.
Publicado: (2025)
por: Loveland, Donald, et al.
Publicado: (2025)
A Pre-training Framework for Relational Data with Information-theoretic Principles
por: Truong, Quang, et al.
Publicado: (2025)
por: Truong, Quang, et al.
Publicado: (2025)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
por: Qin, Zongyue, et al.
Publicado: (2025)
por: Qin, Zongyue, et al.
Publicado: (2025)
Understanding and Scaling Collaborative Filtering Optimization from the Perspective of Matrix Rank
por: Loveland, Donald, et al.
Publicado: (2024)
por: Loveland, Donald, et al.
Publicado: (2024)
FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning
por: Pan, Yijun, et al.
Publicado: (2026)
por: Pan, Yijun, et al.
Publicado: (2026)
FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning
por: Razzaq, Waleed, et al.
Publicado: (2026)
por: Razzaq, Waleed, et al.
Publicado: (2026)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
por: Bai, Xueying, et al.
Publicado: (2024)
por: Bai, Xueying, et al.
Publicado: (2024)
Exploiting ID-Text Complementarity via Ensembling for Sequential Recommendation
por: Collins, Liam, et al.
Publicado: (2025)
por: Collins, Liam, et al.
Publicado: (2025)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
por: Zhu, Jing, et al.
Publicado: (2025)
por: Zhu, Jing, et al.
Publicado: (2025)
Learning Along the Arrow of Time: Hyperbolic Geometry for Backward-Compatible Representation Learning
por: Bui, Ngoc, et al.
Publicado: (2025)
por: Bui, Ngoc, et al.
Publicado: (2025)
One Model for One Graph: A New Perspective for Pretraining with Cross-domain Graphs
por: Liu, Jingzhe, et al.
Publicado: (2024)
por: Liu, Jingzhe, et al.
Publicado: (2024)
Attention Sinks and Outliers in Attention Residuals
por: Luo, Haozheng, et al.
Publicado: (2026)
por: Luo, Haozheng, et al.
Publicado: (2026)
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)
por: Gu, Xiangming, et al.
Publicado: (2024)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
por: Binkowski, Jakub, et al.
Publicado: (2026)
por: Binkowski, Jakub, et al.
Publicado: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
por: Peng, Runyu, et al.
Publicado: (2026)
por: Peng, Runyu, et al.
Publicado: (2026)
ASAP: Attention Sink Anchored Pruning
por: Lee, Jaehyuk, et al.
Publicado: (2026)
por: Lee, Jaehyuk, et al.
Publicado: (2026)
On the Existence and Behavior of Secondary Attention Sinks
por: Wong, Jeffrey T. H., et al.
Publicado: (2025)
por: Wong, Jeffrey T. H., et al.
Publicado: (2025)
Haste Makes Waste: A Simple Approach for Scaling Graph Neural Networks
por: Xue, Rui, et al.
Publicado: (2024)
por: Xue, Rui, et al.
Publicado: (2024)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
por: Son, Seungwoo, et al.
Publicado: (2024)
por: Son, Seungwoo, et al.
Publicado: (2024)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
por: Chen, Yihong, et al.
Publicado: (2026)
por: Chen, Yihong, et al.
Publicado: (2026)
SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models
por: Liu, Junnan, et al.
Publicado: (2026)
por: Liu, Junnan, et al.
Publicado: (2026)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
por: Yu, Zhongzhi, et al.
Publicado: (2024)
por: Yu, Zhongzhi, et al.
Publicado: (2024)
SpecAttn: Speculating Sparse Attention
por: Shah, Harsh
Publicado: (2025)
por: Shah, Harsh
Publicado: (2025)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
por: Willette, Jeffrey, et al.
Publicado: (2025)
por: Willette, Jeffrey, et al.
Publicado: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
por: Fu, Zizhuo, et al.
Publicado: (2026)
por: Fu, Zizhuo, et al.
Publicado: (2026)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
por: Zuhri, Zayd M. K., et al.
Publicado: (2025)
por: Zuhri, Zayd M. K., et al.
Publicado: (2025)
Stochastic Parroting in Temporal Attention -- Regulating the Diagonal Sink
por: Hankemeier, Victoria, et al.
Publicado: (2026)
por: Hankemeier, Victoria, et al.
Publicado: (2026)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
por: Su, Zunhai, et al.
Publicado: (2026)
por: Su, Zunhai, et al.
Publicado: (2026)
Enhancing Item Tokenization for Generative Recommendation through Self-Improvement
por: Chen, Runjin, et al.
Publicado: (2024)
por: Chen, Runjin, et al.
Publicado: (2024)
LLaGA: Large Language and Graph Assistant
por: Chen, Runjin, et al.
Publicado: (2024)
por: Chen, Runjin, et al.
Publicado: (2024)
AutoAL: Automated Active Learning with Differentiable Query Strategy Search
por: Wang, Yifeng, et al.
Publicado: (2024)
por: Wang, Yifeng, et al.
Publicado: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
por: Lee, Heejun, et al.
Publicado: (2023)
por: Lee, Heejun, et al.
Publicado: (2023)
Towards Robust Knowledge Tracing Models via k-Sparse Attention
por: Huang, Shuyan, et al.
Publicado: (2024)
por: Huang, Shuyan, et al.
Publicado: (2024)
Fast Unsupervised Deep Outlier Model Selection with Hypernetworks
por: Ding, Xueying, et al.
Publicado: (2023)
por: Ding, Xueying, et al.
Publicado: (2023)
Scaling Linear Attention with Sparse State Expansion
por: Pan, Yuqi, et al.
Publicado: (2025)
por: Pan, Yuqi, et al.
Publicado: (2025)
Ejemplares similares
-
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
por: Ding, Xueying, et al.
Publicado: (2025) -
How Does Message Passing Improve Collaborative Filtering?
por: Ju, Mingxuan, et al.
Publicado: (2024) -
Robust Training Objectives Improve Embedding-based Retrieval in Industrial Recommendation Systems
por: Kolodner, Matthew, et al.
Publicado: (2024) -
Node Duplication Improves Cold-start Link Prediction
por: Guo, Zhichun, et al.
Publicado: (2024) -
On the Role of Weight Decay in Collaborative Filtering: A Popularity Perspective
por: Loveland, Donald, et al.
Publicado: (2025)