Saved in:
Bibliographic Details
Main Authors: Biton, Dvir David, Friedman, Roy
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.03301
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908864844136448
author Biton, Dvir David
Friedman, Roy
author_facet Biton, Dvir David
Friedman, Roy
contents The rapid adoption of large language models (LLMs) has created demand for faster responses and lower costs. Semantic caching, reusing semantically similar requests via their embeddings, addresses this need but breaks classic cache assumptions and raises new challenges. In this paper, we explore offline policies for semantic caching, proving that implementing an optimal offline policy is NP-hard, and propose several polynomial-time heuristics. We also present online semantic aware cache policies that combine recency, frequency, and locality. Evaluations on diverse datasets show that while frequency based policies are strong baselines, our novel variant improves semantic accuracy. Our findings reveal effective strategies for current systems and highlight substantial headroom for future innovation. All code is open source.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03301
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
Biton, Dvir David
Friedman, Roy
Computation and Language
Artificial Intelligence
Machine Learning
The rapid adoption of large language models (LLMs) has created demand for faster responses and lower costs. Semantic caching, reusing semantically similar requests via their embeddings, addresses this need but breaks classic cache assumptions and raises new challenges. In this paper, we explore offline policies for semantic caching, proving that implementing an optimal offline policy is NP-hard, and propose several polynomial-time heuristics. We also present online semantic aware cache policies that combine recency, frequency, and locality. Evaluations on diverse datasets show that while frequency based policies are strong baselines, our novel variant improves semantic accuracy. Our findings reveal effective strategies for current systems and highlight substantial headroom for future innovation. All code is open source.
title From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.03301