A Generative Caching System for Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Iyengar, Arun, Kundu, Ashish, Kompella, Ramana, Mamidi, Sai Nandan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916659561758720
author Iyengar, Arun
Kundu, Ashish
Kompella, Ramana
Mamidi, Sai Nandan
author_facet Iyengar, Arun
Kundu, Ashish
Kompella, Ramana
Mamidi, Sai Nandan
contents Caching has the potential to be of significant benefit for accessing large language models (LLMs) due to their high latencies which typically range from a small number of seconds to well over a minute. Furthermore, many LLMs charge money for queries; caching thus has a clear monetary benefit. This paper presents a new caching system for improving user experiences with LLMs. In addition to reducing both latencies and monetary costs for accessing LLMs, our system also provides important features that go beyond the performance benefits typically associated with caches. A key feature we provide is generative caching, wherein multiple cached responses can be synthesized to provide answers to queries which have never been seen before. Our generative caches function as repositories of valuable information which can be mined and analyzed. We also improve upon past semantic caching techniques by tailoring the caching algorithms to optimally balance cost and latency reduction with the quality of responses provided. Performance tests indicate that our caches are considerably faster than GPTcache.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Generative Caching System for Large Language Models
Iyengar, Arun
Kundu, Ashish
Kompella, Ramana
Mamidi, Sai Nandan
Databases
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
Caching has the potential to be of significant benefit for accessing large language models (LLMs) due to their high latencies which typically range from a small number of seconds to well over a minute. Furthermore, many LLMs charge money for queries; caching thus has a clear monetary benefit. This paper presents a new caching system for improving user experiences with LLMs. In addition to reducing both latencies and monetary costs for accessing LLMs, our system also provides important features that go beyond the performance benefits typically associated with caches. A key feature we provide is generative caching, wherein multiple cached responses can be synthesized to provide answers to queries which have never been seen before. Our generative caches function as repositories of valuable information which can be mined and analyzed. We also improve upon past semantic caching techniques by tailoring the caching algorithms to optimally balance cost and latency reduction with the quality of responses provided. Performance tests indicate that our caches are considerably faster than GPTcache.
title A Generative Caching System for Large Language Models
topic Databases
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
url https://arxiv.org/abs/2503.17603