XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Weizhuo, Wang, Zhigang, Gu, Yu, Yu, Ge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
di: Behnam, Payman, et al.
Pubblicazione: (2025)
di: Behnam, Payman, et al.
Pubblicazione: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026)
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025)
di: Yu, Bohan, et al.
Pubblicazione: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
di: Li, Kunxi, et al.
Pubblicazione: (2025)
di: Li, Kunxi, et al.
Pubblicazione: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
di: Wang, Luning, et al.
Pubblicazione: (2024)
di: Wang, Luning, et al.
Pubblicazione: (2024)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
di: Nadali, Alireza, et al.
Pubblicazione: (2026)
di: Nadali, Alireza, et al.
Pubblicazione: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
di: Li, Yucheng, et al.
Pubblicazione: (2024)
di: Li, Yucheng, et al.
Pubblicazione: (2024)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
di: Qi, Yanlin, et al.
Pubblicazione: (2026)
di: Qi, Yanlin, et al.
Pubblicazione: (2026)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
di: Liu, Sihao, et al.
Pubblicazione: (2026)
di: Liu, Sihao, et al.
Pubblicazione: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
KV Cache Transform Coding for Compact Storage in LLM Inference
di: Staniszewski, Konrad, et al.
Pubblicazione: (2025)
di: Staniszewski, Konrad, et al.
Pubblicazione: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
di: Ye, Jiancai, et al.
Pubblicazione: (2026)
di: Ye, Jiancai, et al.
Pubblicazione: (2026)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
di: Dong, Zican, et al.
Pubblicazione: (2026)
di: Dong, Zican, et al.
Pubblicazione: (2026)
Inference-Time Hyper-Scaling with KV Cache Compression
di: Łańcucki, Adrian, et al.
Pubblicazione: (2025)
di: Łańcucki, Adrian, et al.
Pubblicazione: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
LongFlow: Efficient KV Cache Compression for Reasoning Models
di: Su, Yi, et al.
Pubblicazione: (2026)
di: Su, Yi, et al.
Pubblicazione: (2026)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
di: Yang, Xintong, et al.
Pubblicazione: (2026)
di: Yang, Xintong, et al.
Pubblicazione: (2026)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
di: Patel, Ishan, et al.
Pubblicazione: (2026)
di: Patel, Ishan, et al.
Pubblicazione: (2026)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
KV Cache Offloading for Context-Intensive Tasks
di: Bocharnikov, Andrey, et al.
Pubblicazione: (2026)
di: Bocharnikov, Andrey, et al.
Pubblicazione: (2026)
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
di: Vejendla, Harshil
Pubblicazione: (2025)
di: Vejendla, Harshil
Pubblicazione: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
di: Yang, Qingyue, et al.
Pubblicazione: (2025)
di: Yang, Qingyue, et al.
Pubblicazione: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
di: Salfati, Samuel
Pubblicazione: (2026)
di: Salfati, Samuel
Pubblicazione: (2026)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
di: Kang, Hao, et al.
Pubblicazione: (2024)
di: Kang, Hao, et al.
Pubblicazione: (2024)
Efficient Long-Context LLM Inference via KV Cache Clustering
di: Hu, Jie, et al.
Pubblicazione: (2025)
di: Hu, Jie, et al.
Pubblicazione: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
di: S, Santhosh G, et al.
Pubblicazione: (2025)
di: S, Santhosh G, et al.
Pubblicazione: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
di: Chang, Chi-Chih, et al.
Pubblicazione: (2025)
di: Chang, Chi-Chih, et al.
Pubblicazione: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
di: Behnam, Payman, et al.
Pubblicazione: (2025) -
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
di: Sun, Hanshi, et al.
Pubblicazione: (2024) -
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026) -
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025) -
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
di: Li, Kunxi, et al.
Pubblicazione: (2025)