DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Guangxuan, Tang, Jiaming, Zuo, Jingwei, Guo, Junxian, Yang, Shang, Tang, Haotian, Fu, Yao, Han, Song |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
Retrieval Head Mechanistically Explains Long-Context Factuality
von: Wu, Wenhao, et al.
Veröffentlicht: (2024)
von: Wu, Wenhao, et al.
Veröffentlicht: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
Optimizing Mixture of Block Attention
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2025)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2025)
XAttention: Block Sparse Attention with Antidiagonal Scoring
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
Efficient Streaming Language Models with Attention Sinks
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2023)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
von: Lin, Ji, et al.
Veröffentlicht: (2023)
von: Lin, Ji, et al.
Veröffentlicht: (2023)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection
von: Tang, Yuying, et al.
Veröffentlicht: (2026)
von: Tang, Yuying, et al.
Veröffentlicht: (2026)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
von: Gu, Zhuohan, et al.
Veröffentlicht: (2024)
von: Gu, Zhuohan, et al.
Veröffentlicht: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
von: Tang, Xiaoya, et al.
Veröffentlicht: (2024)
von: Tang, Xiaoya, et al.
Veröffentlicht: (2024)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
von: Deng, Weishu, et al.
Veröffentlicht: (2025)
von: Deng, Weishu, et al.
Veröffentlicht: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
von: Hu, Qinghao, et al.
Veröffentlicht: (2025)
von: Hu, Qinghao, et al.
Veröffentlicht: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
Analyzing Multi-Head Attention on Trojan BERT Models
von: Wang, Jingwei
Veröffentlicht: (2024)
von: Wang, Jingwei
Veröffentlicht: (2024)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025) -
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024) -
Retrieval Head Mechanistically Explains Long-Context Factuality
von: Wu, Wenhao, et al.
Veröffentlicht: (2024) -
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025) -
Optimizing Mixture of Block Attention
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2025)