Generative Caching for Structurally Similar Prompts and Responses
Fuente:
arXiv
Saved in:
| Main Authors: | Chakraborty, Sarthak, Nath, Suman, Zhang, Xuchao, Bansal, Chetan, Gupta, Indranil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
WebXSkill: Skill Learning for Autonomous Web Agents
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
by: Bansal, Harsh Vardhan
Published: (2025)
by: Bansal, Harsh Vardhan
Published: (2025)
Metric-Fair Prompting: Treating Similar Samples Similarly
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation
by: Bansal, Prakhar, et al.
Published: (2026)
by: Bansal, Prakhar, et al.
Published: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
by: Hao, Jitai, et al.
Published: (2026)
by: Hao, Jitai, et al.
Published: (2026)
SynthAgent: Adapting Web Agents with Synthetic Supervision
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
Anyprefer: An Agentic Framework for Preference Data Synthesis
by: Zhou, Yiyang, et al.
Published: (2025)
by: Zhou, Yiyang, et al.
Published: (2025)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
by: Gim, In, et al.
Published: (2023)
by: Gim, In, et al.
Published: (2023)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
LLM driven Text-to-Table Generation through Sub-Tasks Guidance and Iterative Refinement
by: C, Rajmohan, et al.
Published: (2025)
by: C, Rajmohan, et al.
Published: (2025)
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
by: Lee, Kun-Hui, et al.
Published: (2025)
by: Lee, Kun-Hui, et al.
Published: (2025)
Extreme Speech Classification in the Era of LLMs: Exploring Open-Source and Proprietary Models
by: Mahajan, Sarthak, et al.
Published: (2025)
by: Mahajan, Sarthak, et al.
Published: (2025)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
by: Zhang, Ruochen, et al.
Published: (2024)
by: Zhang, Ruochen, et al.
Published: (2024)
Pruning as a Defense: Reducing Memorization in Large Language Models
by: Gupta, Mansi, et al.
Published: (2025)
by: Gupta, Mansi, et al.
Published: (2025)
Layers at Similar Depths Generate Similar Activations Across LLM Architectures
by: Wolfram, Christopher, et al.
Published: (2025)
by: Wolfram, Christopher, et al.
Published: (2025)
Examining the Mental Health Impact of Misinformation on Social Media Using a Hybrid Transformer-Based Approach
by: Arora, Sarvesh, et al.
Published: (2025)
by: Arora, Sarvesh, et al.
Published: (2025)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
by: Kucia, Filip J., et al.
Published: (2026)
by: Kucia, Filip J., et al.
Published: (2026)
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Meaning Typed Prompting: A Technique for Efficient, Reliable Structured Output Generation
by: Irugalbandara, Chandra
Published: (2024)
by: Irugalbandara, Chandra
Published: (2024)
Auto-Demo Prompting: Leveraging Generated Outputs as Demonstrations for Enhanced Batch Prompting
by: Feng, Longyu, et al.
Published: (2024)
by: Feng, Longyu, et al.
Published: (2024)
From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors
by: Zhou, Yitian, et al.
Published: (2026)
by: Zhou, Yitian, et al.
Published: (2026)
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Generating Gender Alternatives in Machine Translation
by: Garg, Sarthak, et al.
Published: (2024)
by: Garg, Sarthak, et al.
Published: (2024)
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
by: Kjorvezir, Denica, et al.
Published: (2026)
by: Kjorvezir, Denica, et al.
Published: (2026)
No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
by: Rajgaria, Abhishek, et al.
Published: (2025)
by: Rajgaria, Abhishek, et al.
Published: (2025)
Building Domain-Specific Small Language Models via Guided Data Generation
by: Kumar, Aman, et al.
Published: (2025)
by: Kumar, Aman, et al.
Published: (2025)
TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL
by: Su, Jinbo, et al.
Published: (2026)
by: Su, Jinbo, et al.
Published: (2026)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
by: Wu, Chenyuan, et al.
Published: (2024)
by: Wu, Chenyuan, et al.
Published: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Budget-Aware Agentic Routing via Boundary-Guided Training
by: Zhang, Caiqi, et al.
Published: (2026)
by: Zhang, Caiqi, et al.
Published: (2026)
Generative Prompt Internalization
by: Shin, Haebin, et al.
Published: (2024)
by: Shin, Haebin, et al.
Published: (2024)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
by: Zhao, Xinye, et al.
Published: (2025)
by: Zhao, Xinye, et al.
Published: (2025)
Automatic Prompt Generation via Adaptive Selection of Prompting Techniques
by: Ikenoue, Yohei, et al.
Published: (2025)
by: Ikenoue, Yohei, et al.
Published: (2025)
SCOPE: A Generative Approach for LLM Prompt Compression
by: Zhang, Tinghui, et al.
Published: (2025)
by: Zhang, Tinghui, et al.
Published: (2025)
InsightNet: Structured Insight Mining from Customer Feedback
by: Mukku, Sandeep Sricharan, et al.
Published: (2024)
by: Mukku, Sandeep Sricharan, et al.
Published: (2024)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
by: Saxena, Yash, et al.
Published: (2024)
by: Saxena, Yash, et al.
Published: (2024)
Similar Items
-
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025) -
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024) -
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
by: Gupta, Taneesh, et al.
Published: (2024) -
WebXSkill: Skill Learning for Autonomous Web Agents
by: Wang, Zhaoyang, et al.
Published: (2026) -
LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
by: Bansal, Harsh Vardhan
Published: (2025)