Efficient Streaming Language Models with Attention Sinks
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Guangxuan, Tian, Yuandong, Chen, Beidi, Han, Song, Lewis, Mike |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023)
by: Tian, Yuandong, et al.
Published: (2023)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
Attention Sinks in Diffusion Language Models
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)
by: Xiao, Guangxuan, et al.
Published: (2022)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
When Attention Sink Emerges in Language Models: An Empirical View
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
Spectral Filters, Dark Signals, and Attention Sinks
by: Cancedda, Nicola
Published: (2024)
by: Cancedda, Nicola
Published: (2024)
Positional Encoding via Token-Aware Phase Attention
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
by: Yang, Kevin, et al.
Published: (2023)
by: Yang, Kevin, et al.
Published: (2023)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
by: Pan, Dayan, et al.
Published: (2025)
by: Pan, Dayan, et al.
Published: (2025)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2025)
by: Luo, Jiayun, et al.
Published: (2025)
On the Surprising Effectiveness of Attention Transfer for Vision Transformers
by: Li, Alexander C., et al.
Published: (2024)
by: Li, Alexander C., et al.
Published: (2024)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
by: Zheng, Haizhong, et al.
Published: (2024)
by: Zheng, Haizhong, et al.
Published: (2024)
Assessing Large Language Models for Online Extremism Research: Identification, Explanation, and New Knowledge
by: Dong, Beidi, et al.
Published: (2024)
by: Dong, Beidi, et al.
Published: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
You Only Use Reactive Attention Slice For Long Context Retrieval
by: Soh, Yun Joon, et al.
Published: (2024)
by: Soh, Yun Joon, et al.
Published: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
One Token Is Enough: Improving Diffusion Language Models with a Sink Token
by: Zhang, Zihou, et al.
Published: (2026)
by: Zhang, Zihou, et al.
Published: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Attention-Aligned Reasoning for Large Language Models
by: Zhang, Hongxiang, et al.
Published: (2025)
by: Zhang, Hongxiang, et al.
Published: (2025)
Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
by: Tian, Yuandong
Published: (2024)
by: Tian, Yuandong
Published: (2024)
Param$Δ$ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost
by: Cao, Sheng, et al.
Published: (2025)
by: Cao, Sheng, et al.
Published: (2025)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
by: Shan, Hongyu, et al.
Published: (2025)
by: Shan, Hongyu, et al.
Published: (2025)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Efficient Attention Mechanisms for Large Language Models: A Survey
by: Sun, Yutao, et al.
Published: (2025)
by: Sun, Yutao, et al.
Published: (2025)
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Sliding Window Attention Training for Efficient Large Language Models
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
by: Grand, Gabriel, et al.
Published: (2024)
by: Grand, Gabriel, et al.
Published: (2024)
Self-Selected Attention Span for Accelerating Large Language Model Inference
by: Jin, Tian, et al.
Published: (2024)
by: Jin, Tian, et al.
Published: (2024)
StreamAdapter: Efficient Test Time Adaptation from Contextual Streams
by: Muhtar, Dilxat, et al.
Published: (2024)
by: Muhtar, Dilxat, et al.
Published: (2024)
Similar Items
-
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023) -
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
by: Zhou, Yang, et al.
Published: (2025) -
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025) -
Attention Sinks in Diffusion Language Models
by: Rulli, Maximo Eduardo, et al.
Published: (2025) -
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)