Generative Caching for Structurally Similar Prompts and Responses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chakraborty, Sarthak, Nath, Suman, Zhang, Xuchao, Bansal, Chetan, Gupta, Indranil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912724113424384
author Chakraborty, Sarthak
Nath, Suman
Zhang, Xuchao
Bansal, Chetan
Gupta, Indranil
author_facet Chakraborty, Sarthak
Nath, Suman
Zhang, Xuchao
Bansal, Chetan
Gupta, Indranil
contents Large Language Models (LLMs) are increasingly being used to plan, reason, and execute tasks across diverse scenarios. In use cases like repeatable workflows and agentic settings, prompts are often reused with minor variations while having a similar structure for recurring tasks. This opens up opportunities for caching. However, exact prompt matching fails on such structurally similar prompts, while semantic caching may produce incorrect responses by ignoring critical differences. To address this, we introduce \ourmethod{}, a generative cache that produces variation-aware responses for structurally similar prompts. \ourmethod{} identifies reusable response patterns across similar prompt structures and synthesizes customized outputs for new requests. We show that \ourmethod{} achieves 83\% cache hit rate, while having minimal incorrect hits on datasets without prompt repetition. In agentic workflows, it improves cache hit rate by $\sim$20\% and reduces end-to-end execution latency by $\sim$34\% compared to standard prompt matching.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Caching for Structurally Similar Prompts and Responses
Chakraborty, Sarthak
Nath, Suman
Zhang, Xuchao
Bansal, Chetan
Gupta, Indranil
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) are increasingly being used to plan, reason, and execute tasks across diverse scenarios. In use cases like repeatable workflows and agentic settings, prompts are often reused with minor variations while having a similar structure for recurring tasks. This opens up opportunities for caching. However, exact prompt matching fails on such structurally similar prompts, while semantic caching may produce incorrect responses by ignoring critical differences. To address this, we introduce \ourmethod{}, a generative cache that produces variation-aware responses for structurally similar prompts. \ourmethod{} identifies reusable response patterns across similar prompt structures and synthesizes customized outputs for new requests. We show that \ourmethod{} achieves 83\% cache hit rate, while having minimal incorrect hits on datasets without prompt repetition. In agentic workflows, it improves cache hit rate by $\sim$20\% and reduces end-to-end execution latency by $\sim$34\% compared to standard prompt matching.
title Generative Caching for Structurally Similar Prompts and Responses
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.17565