S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Xihui, Zhang, Yunan, Ge, Suyu, Ren, Liliang, Patra, Barun, Chaudhary, Vishrav, Peng, Hao, Song, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024)
by: Ge, Suyu, et al.
Published: (2024)
A Practical Analysis of Human Alignment with *PO
by: Ahrabian, Kian, et al.
Published: (2024)
by: Ahrabian, Kian, et al.
Published: (2024)
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023)
by: Monea, Giovanni, et al.
Published: (2023)
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)
by: Bjorck, Johan, et al.
Published: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
by: Hua, Kai, et al.
Published: (2025)
by: Hua, Kai, et al.
Published: (2025)
Attention Mechanism and Heuristic Approach: Context-Aware File Ranking Using Multi-Head Self-Attention
by: Sharma, Pradeep Kumar, et al.
Published: (2026)
by: Sharma, Pradeep Kumar, et al.
Published: (2026)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization
by: Karaman, Batuhan K., et al.
Published: (2024)
by: Karaman, Batuhan K., et al.
Published: (2024)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
by: Ge, Suyu, et al.
Published: (2023)
by: Ge, Suyu, et al.
Published: (2023)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
by: Ahuja, Sanchit, et al.
Published: (2025)
by: Ahuja, Sanchit, et al.
Published: (2025)
Knocking-Heads Attention
by: Zhou, Zhanchao, et al.
Published: (2025)
by: Zhou, Zhanchao, et al.
Published: (2025)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
SAS: Simulated Attention Score
by: Zheng, Chuanyang, et al.
Published: (2025)
by: Zheng, Chuanyang, et al.
Published: (2025)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
by: Ahrabian, Kian, et al.
Published: (2024)
by: Ahrabian, Kian, et al.
Published: (2024)
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
Long Context Pre-Training with Lighthouse Attention
by: Peng, Bowen, et al.
Published: (2026)
by: Peng, Bowen, et al.
Published: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Attention Heads of Large Language Models: A Survey
by: Zheng, Zifan, et al.
Published: (2024)
by: Zheng, Zifan, et al.
Published: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
by: Ghaddar, Abbas, et al.
Published: (2026)
by: Ghaddar, Abbas, et al.
Published: (2026)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
by: Ai, Xuan, et al.
Published: (2026)
by: Ai, Xuan, et al.
Published: (2026)
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025)
by: Zadouri, Ted, et al.
Published: (2025)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
by: Li, Zichong, et al.
Published: (2026)
by: Li, Zichong, et al.
Published: (2026)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
by: Liu, Siran, et al.
Published: (2026)
by: Liu, Siran, et al.
Published: (2026)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
by: Liu, Xiaoran, et al.
Published: (2024)
by: Liu, Xiaoran, et al.
Published: (2024)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
Similar Items
-
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024) -
A Practical Analysis of Human Alignment with *PO
by: Ahrabian, Kian, et al.
Published: (2024) -
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024) -
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023) -
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)