ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Lu, Tao, Ze, Huang, Yong, Li, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
Sparse Attention across Multiple-context KV Cache
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
EntmaxKV: Support-Aware Decoding for Entmax Attention
von: Duarte, Gonçalo, et al.
Veröffentlicht: (2026)
von: Duarte, Gonçalo, et al.
Veröffentlicht: (2026)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
von: Guo, Zhiyu, et al.
Veröffentlicht: (2024)
von: Guo, Zhiyu, et al.
Veröffentlicht: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
LongFlow: Efficient KV Cache Compression for Reasoning Models
von: Su, Yi, et al.
Veröffentlicht: (2026)
von: Su, Yi, et al.
Veröffentlicht: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
von: Wong, Liang Ze
Veröffentlicht: (2025)
von: Wong, Liang Ze
Veröffentlicht: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
von: Gao, Bin, et al.
Veröffentlicht: (2024)
von: Gao, Bin, et al.
Veröffentlicht: (2024)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
von: Song, Guanghui, et al.
Veröffentlicht: (2025)
von: Song, Guanghui, et al.
Veröffentlicht: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
von: Sharma, Akshat, et al.
Veröffentlicht: (2024)
von: Sharma, Akshat, et al.
Veröffentlicht: (2024)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024) -
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024) -
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025) -
Sparse Attention across Multiple-context KV Cache
von: Cao, Ziyi, et al.
Veröffentlicht: (2025) -
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)