Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jungwoo, Kim, Minsang, Lee, Jaeheon, Moon, Chanwoo, Kim, Heejin, Hwang, Taeho, Chung, Woosuk, Kim, Yeseong, Lee, Sungjin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025)
by: Kim, Jang-Hyun, et al.
Published: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
RELOAD: A Robust and Efficient Learned Query Optimizer for Database Systems
by: Lee, Seokwon, et al.
Published: (2026)
by: Lee, Seokwon, et al.
Published: (2026)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
by: Zeng, Pai, et al.
Published: (2024)
by: Zeng, Pai, et al.
Published: (2024)
Exploring Cross-Stage Adversarial Transferability in Class-Incremental Continual Learning
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
Diffusion Bridge AutoEncoders for Unsupervised Representation Learning
by: Kim, Yeongmin, et al.
Published: (2024)
by: Kim, Yeongmin, et al.
Published: (2024)
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
by: Hu, Zhengding, et al.
Published: (2025)
by: Hu, Zhengding, et al.
Published: (2025)
Measuring Sample Importance in Data Pruning for Language Models based on Information Entropy
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Hierarchical Position Embedding of Graphs with Landmarks and Clustering for Link Prediction
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Accelerating String-Key Learned Index Structures via Memoization-based Incremental Training
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
by: Kim, Yeongmin, et al.
Published: (2026)
by: Kim, Yeongmin, et al.
Published: (2026)
RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
by: Wang, Zhengchao, et al.
Published: (2025)
by: Wang, Zhengchao, et al.
Published: (2025)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Hydro: Adaptive Query Processing of ML Queries
by: Kakkar, Gaurav Tarlok, et al.
Published: (2024)
by: Kakkar, Gaurav Tarlok, et al.
Published: (2024)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
by: An, Selim, et al.
Published: (2026)
by: An, Selim, et al.
Published: (2026)
EdgeServe: A Streaming System for Decentralized Model Serving
by: Shaowang, Ted, et al.
Published: (2023)
by: Shaowang, Ted, et al.
Published: (2023)
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
by: Wu, Fangzhou, et al.
Published: (2025)
by: Wu, Fangzhou, et al.
Published: (2025)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
by: Han, Seungyub, et al.
Published: (2026)
by: Han, Seungyub, et al.
Published: (2026)
Scalable bayesian shadow tomography for quantum property estimation with set transformers
by: Cha, Hyunho, et al.
Published: (2025)
by: Cha, Hyunho, et al.
Published: (2025)
NEXT-EVAL: Next Evaluation of Traditional and LLM Web Data Record Extraction
by: Kim, Soyeon, et al.
Published: (2025)
by: Kim, Soyeon, et al.
Published: (2025)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
by: Hwang, Soon, et al.
Published: (2025)
by: Hwang, Soon, et al.
Published: (2025)
PnPXAI: A Universal XAI Framework Providing Automatic Explanations Across Diverse Modalities and Models
by: Kim, Seongun, et al.
Published: (2025)
by: Kim, Seongun, et al.
Published: (2025)
Effective Dataset Distillation for Spatio-Temporal Forecasting with Bi-dimensional Compression
by: Kwon, Taehyung, et al.
Published: (2026)
by: Kwon, Taehyung, et al.
Published: (2026)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
by: Liu, Banruo, et al.
Published: (2025)
by: Liu, Banruo, et al.
Published: (2025)
BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
by: Kim, Mingi, et al.
Published: (2026)
by: Kim, Mingi, et al.
Published: (2026)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
by: Kim, Minsang, et al.
Published: (2026)
by: Kim, Minsang, et al.
Published: (2026)
TranSQL+: Serving Large Language Models with SQL on Low-Resource Hardware
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
by: Bergman, Shai, et al.
Published: (2025)
by: Bergman, Shai, et al.
Published: (2025)
Dynamic data summarization for hierarchical spatial clustering
by: Abduaziz, Kayumov, et al.
Published: (2024)
by: Abduaziz, Kayumov, et al.
Published: (2024)
Selection of the Most Probable Best
by: Kim, Taeho, et al.
Published: (2022)
by: Kim, Taeho, et al.
Published: (2022)
Accelerating Storage-Based Training for Graph Neural Networks
by: Jang, Myung-Hwan, et al.
Published: (2026)
by: Jang, Myung-Hwan, et al.
Published: (2026)
M2: An Analytic System with Specialized Storage Engines for Multi-Model Workloads
by: Koo, Kyoseung, et al.
Published: (2025)
by: Koo, Kyoseung, et al.
Published: (2025)
Optimizing Input Data Collection for Ranking and Selection
by: Song, Eunhye, et al.
Published: (2025)
by: Song, Eunhye, et al.
Published: (2025)
Data-Driven Sequential Sampling for Tail Risk Mitigation
by: Ahn, Dohyun, et al.
Published: (2025)
by: Ahn, Dohyun, et al.
Published: (2025)
Hype or Heuristic? Quantum Reinforcement Learning for Join Order Optimisation
by: Franz, Maja, et al.
Published: (2024)
by: Franz, Maja, et al.
Published: (2024)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
An LLM-Guided Query-Aware Inference System for GNN Models on Large Knowledge Graphs
by: Afandi, Waleed, et al.
Published: (2026)
by: Afandi, Waleed, et al.
Published: (2026)
Graph Spectral Filtering with Chebyshev Interpolation for Recommendation
by: Kim, Chanwoo, et al.
Published: (2025)
by: Kim, Chanwoo, et al.
Published: (2025)
Similar Items
-
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
by: Kim, Jungwoo, et al.
Published: (2025) -
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025) -
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
by: Kim, Kihyun, et al.
Published: (2025) -
RELOAD: A Robust and Efficient Learned Query Optimizer for Database Systems
by: Lee, Seokwon, et al.
Published: (2026) -
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)