RAC: Relation-Aware Cache Replacement for Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Yuchong, Xu, Zihuan, Ni, Wangze, Cheng, Peng, Chen, Lei, Lin, Xuemin, Shen, Heng Tao, Ren, Kui
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917293778272256
author Wu, Yuchong
Xu, Zihuan
Ni, Wangze
Cheng, Peng
Chen, Lei
Lin, Xuemin
Shen, Heng Tao
Ren, Kui
author_facet Wu, Yuchong
Xu, Zihuan
Ni, Wangze
Cheng, Peng
Chen, Lei
Lin, Xuemin
Shen, Heng Tao
Ren, Kui
contents The scaling of Large Language Model (LLM) services faces significant cost and latency challenges, making effective caching under tight capacity crucial. Existing cache replacement policies, from heuristics to learning-based methods, predominantly rely on limited-window statistics such as recency and frequency. We show these signals are not robust for real-world LLM workloads, which exhibit long reuse distances and sparse local recurrence. To address these limitations, we propose Relation-Aware Cache (RAC), an online eviction strategy that leverages semantic relations among requests to guide eviction decisions. RAC synthesizes two relation-aware signals: (1) Topical Prevalence, which aggregates access evidence at the topic level to capture long-horizon reuse; and (2) Structural Importance, which leverages local intra-topic dependency structure to discriminate entries by their future reuse value. Extensive evaluations show that RAC maintains high effectiveness across diverse workloads, consistently surpassing state-of-the-art baselines by 20%--30% in cache hit ratio.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21547
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RAC: Relation-Aware Cache Replacement for Large Language Models
Wu, Yuchong
Xu, Zihuan
Ni, Wangze
Cheng, Peng
Chen, Lei
Lin, Xuemin
Shen, Heng Tao
Ren, Kui
Databases
The scaling of Large Language Model (LLM) services faces significant cost and latency challenges, making effective caching under tight capacity crucial. Existing cache replacement policies, from heuristics to learning-based methods, predominantly rely on limited-window statistics such as recency and frequency. We show these signals are not robust for real-world LLM workloads, which exhibit long reuse distances and sparse local recurrence. To address these limitations, we propose Relation-Aware Cache (RAC), an online eviction strategy that leverages semantic relations among requests to guide eviction decisions. RAC synthesizes two relation-aware signals: (1) Topical Prevalence, which aggregates access evidence at the topic level to capture long-horizon reuse; and (2) Structural Importance, which leverages local intra-topic dependency structure to discriminate entries by their future reuse value. Extensive evaluations show that RAC maintains high effectiveness across diverse workloads, consistently surpassing state-of-the-art baselines by 20%--30% in cache hit ratio.
title RAC: Relation-Aware Cache Replacement for Large Language Models
topic Databases
url https://arxiv.org/abs/2602.21547