Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916918111240192 |
|---|---|
| author | Kim, Jungwoo Kim, Minsang Lee, Jaeheon Moon, Chanwoo Kim, Heejin Hwang, Taeho Chung, Woosuk Kim, Yeseong Lee, Sungjin |
| author_facet | Kim, Jungwoo Kim, Minsang Lee, Jaeheon Moon, Chanwoo Kim, Heejin Hwang, Taeho Chung, Woosuk Kim, Yeseong Lee, Sungjin |
| contents | Serving Large Language Models (LLMs) at scale requires meeting strict Service Level Objectives (SLOs) under severe computational and memory constraints. Nevertheless, traditional caching strategies fall short: exact-matching and prefix caches neglect query semantics, while state-of-the-art semantic caches remain confined to traditional intuitions, offering little conceptual departure. Building on this, we present SISO, a semantic caching system that redefines efficiency for LLM serving. SISO introduces centroid-based caching to maximize coverage with minimal memory, locality-aware replacement to preserve high-value entries, and dynamic thresholding to balance accuracy and latency under varying workloads. Across diverse datasets, SISO delivers up to 1.71$\times$ higher hit ratios and consistently stronger SLO attainment compared to state-of-the-art systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_18736 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics Kim, Jungwoo Kim, Minsang Lee, Jaeheon Moon, Chanwoo Kim, Heejin Hwang, Taeho Chung, Woosuk Kim, Yeseong Lee, Sungjin Databases Machine Learning Serving Large Language Models (LLMs) at scale requires meeting strict Service Level Objectives (SLOs) under severe computational and memory constraints. Nevertheless, traditional caching strategies fall short: exact-matching and prefix caches neglect query semantics, while state-of-the-art semantic caches remain confined to traditional intuitions, offering little conceptual departure. Building on this, we present SISO, a semantic caching system that redefines efficiency for LLM serving. SISO introduces centroid-based caching to maximize coverage with minimal memory, locality-aware replacement to preserve high-value entries, and dynamic thresholding to balance accuracy and latency under varying workloads. Across diverse datasets, SISO delivers up to 1.71$\times$ higher hit ratios and consistently stronger SLO attainment compared to state-of-the-art systems. |
| title | Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics |
| topic | Databases Machine Learning |
| url | https://arxiv.org/abs/2508.18736 |