EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Minsoo, Kundu, Arnav, Kim, Han-Byul, Dixit, Richa, Cho, Minsik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026)
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
dKV-Cache: The Cache for Diffusion Language Models
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
RefreshKV: Updating Small KV Cache During Long-form Generation
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
TIDE: Every Layer Knows the Token Beneath the Context
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2026)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2026)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
von: Peng, Junjie, et al.
Veröffentlicht: (2026)
von: Peng, Junjie, et al.
Veröffentlicht: (2026)
LongFlow: Efficient KV Cache Compression for Reasoning Models
von: Su, Yi, et al.
Veröffentlicht: (2026)
von: Su, Yi, et al.
Veröffentlicht: (2026)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
von: Lin, Jian, et al.
Veröffentlicht: (2026)
von: Lin, Jian, et al.
Veröffentlicht: (2026)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
von: Qi, Yanlin, et al.
Veröffentlicht: (2026)
von: Qi, Yanlin, et al.
Veröffentlicht: (2026)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
von: Liu, Tengxuan, et al.
Veröffentlicht: (2025)
von: Liu, Tengxuan, et al.
Veröffentlicht: (2025)
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
von: Lu, Kuan, et al.
Veröffentlicht: (2025)
von: Lu, Kuan, et al.
Veröffentlicht: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
von: Cho, Minsik, et al.
Veröffentlicht: (2024) -
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026) -
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025) -
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026) -
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)