ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhuorui, Zhang, Chen, Song, Dawei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
by: Zuo, Youhui, et al.
Published: (2025)
by: Zuo, Youhui, et al.
Published: (2025)
Beyond the Speculative Game: A Survey of Speculative Execution in Large Language Models
by: Zhang, Chen, et al.
Published: (2024)
by: Zhang, Chen, et al.
Published: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
by: Fei, Weizhi, et al.
Published: (2025)
by: Fei, Weizhi, et al.
Published: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
by: Ai, Xuan, et al.
Published: (2026)
by: Ai, Xuan, et al.
Published: (2026)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
by: Liu, Runheng, et al.
Published: (2024)
by: Liu, Runheng, et al.
Published: (2024)
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
by: Zhang, Wuwei, et al.
Published: (2025)
by: Zhang, Wuwei, et al.
Published: (2025)
Retrieval Head Mechanistically Explains Long-Context Factuality
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
by: Liu, Siran, et al.
Published: (2026)
by: Liu, Siran, et al.
Published: (2026)
LongEmbed: Extending Embedding Models for Long Context Retrieval
by: Zhu, Dawei, et al.
Published: (2024)
by: Zhu, Dawei, et al.
Published: (2024)
Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
by: Chen, Jiamin, et al.
Published: (2025)
by: Chen, Jiamin, et al.
Published: (2025)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
ParetoRAG: Leveraging Sentence-Context Attention for Robust and Efficient Retrieval-Augmented Generation
by: Yao, Ruobing, et al.
Published: (2025)
by: Yao, Ruobing, et al.
Published: (2025)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2025)
by: You, Zeng, et al.
Published: (2025)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
by: Zhang, Yuwei, et al.
Published: (2025)
by: Zhang, Yuwei, et al.
Published: (2025)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
Inference Scaling for Long-Context Retrieval Augmented Generation
by: Yue, Zhenrui, et al.
Published: (2024)
by: Yue, Zhenrui, et al.
Published: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
by: Ye, Xiaoju, et al.
Published: (2025)
by: Ye, Xiaoju, et al.
Published: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
Exclusive Self Attention
by: Zhai, Shuangfei
Published: (2026)
by: Zhai, Shuangfei
Published: (2026)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
by: Hu, Xiang, et al.
Published: (2024)
by: Hu, Xiang, et al.
Published: (2024)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024)
by: Ge, Suyu, et al.
Published: (2024)
Efficient Streaming Language Models with Attention Sinks
by: Xiao, Guangxuan, et al.
Published: (2023)
by: Xiao, Guangxuan, et al.
Published: (2023)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
Similar Items
-
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024) -
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
by: Zuo, Youhui, et al.
Published: (2025) -
Beyond the Speculative Game: A Survey of Speculative Execution in Large Language Models
by: Zhang, Chen, et al.
Published: (2024) -
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024) -
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
by: Fei, Weizhi, et al.
Published: (2025)