Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyungwoo, Kim, Kihyun, Kim, Jinwoo, So, Jungmin, Cha, Myung-Hoon, Kim, Hong-Yeon, Kim, James J., Kim, Youngjae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
by: Kim, Yoochan, et al.
Published: (2024)
by: Kim, Yoochan, et al.
Published: (2024)
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Investigating the Integrated Digital Interventions Delivered by a Therapeutic Companion Agent for Young Adults with Symptoms of Depression: A Proof-of-Concept Study
by: Yoo, Youngjae, et al.
Published: (2025)
by: Yoo, Youngjae, et al.
Published: (2025)
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
by: Jeon, Hyesung, et al.
Published: (2026)
by: Jeon, Hyesung, et al.
Published: (2026)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Rigidity of smooth finite-time blow-up for equivariant self-dual Chern-Simons-Schrödinger equation
by: Kim, Kihyun
Published: (2022)
by: Kim, Kihyun
Published: (2022)
Electrochemical characteristics of Si/Mo multilayer anode for Li ion batteries
by: Myung-Hoon Kim
Published: (2007)
by: Myung-Hoon Kim
Published: (2007)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
Metal‐Free Submicron‐Hollow‐Fiber Conjugated Polymer Sponges for Efficient Pollutant Removal and Thermal Insulation
by: Songah Jeong, et al.
Published: (2026)
by: Songah Jeong, et al.
Published: (2026)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
Multidimensional Analysis of Visitor Interaction With Museum Apps: Designing for Enhanced Museum Experiences
by: Jihyun Kim, et al.
Published: (2025)
by: Jihyun Kim, et al.
Published: (2025)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
by: Ahn, Jinwoo, et al.
Published: (2026)
by: Ahn, Jinwoo, et al.
Published: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
by: Cho, Youngjae, et al.
Published: (2026)
by: Cho, Youngjae, et al.
Published: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
by: Yun, Jungmin, et al.
Published: (2024)
by: Yun, Jungmin, et al.
Published: (2024)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
BERT-APC: A Reference-free Framework for Automatic Pitch Correction via Musical Context Inference
by: Kim, Sungjae, et al.
Published: (2025)
by: Kim, Sungjae, et al.
Published: (2025)
Simulation-Free Training of Neural ODEs on Paired Data
by: Kim, Semin, et al.
Published: (2024)
by: Kim, Semin, et al.
Published: (2024)
Light-Wave Engineering for Selective Polarization of a Single $\mathbf{Q}$ Valley in Transition Metal Dichalcogenides
by: Kim, Youngjae
Published: (2025)
by: Kim, Youngjae
Published: (2025)
Pseudospins revealed through the giant dynamical Franz-Keldysh effect in massless Dirac materials
by: Kim, Youngjae
Published: (2024)
by: Kim, Youngjae
Published: (2024)
Common Ownership and Auditor Sharing
by: Young Hoon Kim
Published: (2026)
by: Young Hoon Kim
Published: (2026)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025)
by: Kim, Hee-Seon, et al.
Published: (2025)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025)
by: Kim, Jang-Hyun, et al.
Published: (2025)
Construction of smooth chiral finite-time blow-up solutions to Calogero--Moser derivative nonlinear Schrödinger equation
by: Kim, Kihyun, et al.
Published: (2024)
by: Kim, Kihyun, et al.
Published: (2024)
Self-Abstraction Learning for Effective and Stable Training of Deep Neural Networks
by: Cho, Wonyong, et al.
Published: (2026)
by: Cho, Wonyong, et al.
Published: (2026)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
by: Kim, Munsik
Published: (2026)
by: Kim, Munsik
Published: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
Beyond KV Caching: Shared Attention for Efficient LLMs
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
On classification of global dynamics for energy-critical equivariant harmonic map heat flows and radial nonlinear heat equation
by: Kim, Kihyun, et al.
Published: (2024)
by: Kim, Kihyun, et al.
Published: (2024)
Rigidity results in multi-bubble dynamics for non-radial energy-critical heat equation
by: Kim, Kihyun, et al.
Published: (2026)
by: Kim, Kihyun, et al.
Published: (2026)
Construction of infinite time bubble tower solutions to critical wave maps equation
by: Hwang, Seunghwan, et al.
Published: (2026)
by: Hwang, Seunghwan, et al.
Published: (2026)
On classification of global dynamics for energy‐critical equivariant harmonic map heat flows and radial nonlinear heat equation
by: Kihyun Kim, et al.
Published: (2025)
by: Kihyun Kim, et al.
Published: (2025)
Program Synthesis is $Σ_3^0$-Complete
by: Kim, Jinwoo
Published: (2024)
by: Kim, Jinwoo
Published: (2024)
A Review of Image Retrieval Techniques: Data Augmentation and Adversarial Learning Approaches
by: Jinwoo, Kim
Published: (2024)
by: Jinwoo, Kim
Published: (2024)
Similar Items
-
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
by: Kim, Kihyun, et al.
Published: (2025) -
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
by: Kim, Yoochan, et al.
Published: (2024) -
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025) -
Investigating the Integrated Digital Interventions Delivered by a Therapeutic Companion Agent for Young Adults with Symptoms of Depression: A Proof-of-Concept Study
by: Yoo, Youngjae, et al.
Published: (2025) -
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
by: Jeon, Hyesung, et al.
Published: (2026)