Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Jungwoo, Kim, Minsang, Lee, Jaeheon, Moon, Chanwoo, Kim, Heejin, Hwang, Taeho, Chung, Woosuk, Kim, Yeseong, Lee, Sungjin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916918111240192
author Kim, Jungwoo
Kim, Minsang
Lee, Jaeheon
Moon, Chanwoo
Kim, Heejin
Hwang, Taeho
Chung, Woosuk
Kim, Yeseong
Lee, Sungjin
author_facet Kim, Jungwoo
Kim, Minsang
Lee, Jaeheon
Moon, Chanwoo
Kim, Heejin
Hwang, Taeho
Chung, Woosuk
Kim, Yeseong
Lee, Sungjin
contents Serving Large Language Models (LLMs) at scale requires meeting strict Service Level Objectives (SLOs) under severe computational and memory constraints. Nevertheless, traditional caching strategies fall short: exact-matching and prefix caches neglect query semantics, while state-of-the-art semantic caches remain confined to traditional intuitions, offering little conceptual departure. Building on this, we present SISO, a semantic caching system that redefines efficiency for LLM serving. SISO introduces centroid-based caching to maximize coverage with minimal memory, locality-aware replacement to preserve high-value entries, and dynamic thresholding to balance accuracy and latency under varying workloads. Across diverse datasets, SISO delivers up to 1.71$\times$ higher hit ratios and consistently stronger SLO attainment compared to state-of-the-art systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18736
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
Kim, Jungwoo
Kim, Minsang
Lee, Jaeheon
Moon, Chanwoo
Kim, Heejin
Hwang, Taeho
Chung, Woosuk
Kim, Yeseong
Lee, Sungjin
Databases
Machine Learning
Serving Large Language Models (LLMs) at scale requires meeting strict Service Level Objectives (SLOs) under severe computational and memory constraints. Nevertheless, traditional caching strategies fall short: exact-matching and prefix caches neglect query semantics, while state-of-the-art semantic caches remain confined to traditional intuitions, offering little conceptual departure. Building on this, we present SISO, a semantic caching system that redefines efficiency for LLM serving. SISO introduces centroid-based caching to maximize coverage with minimal memory, locality-aware replacement to preserve high-value entries, and dynamic thresholding to balance accuracy and latency under varying workloads. Across diverse datasets, SISO delivers up to 1.71$\times$ higher hit ratios and consistently stronger SLO attainment compared to state-of-the-art systems.
title Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
topic Databases
Machine Learning
url https://arxiv.org/abs/2508.18736