Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Bergman, Shai, Kermarrec, Anne-Marie, Petrescu, Diana, Pires, Rafael, Randl, Mathis, de Vos, Martijn, Zhang, Ji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
Catapults to the Rescue: Accelerating Vector Search by Exploiting Query Locality
by: Abuzakuk, Sami, et al.
Published: (2026)
by: Abuzakuk, Sami, et al.
Published: (2026)
Effective LoRA Adapter Routing using Task Representations
by: Dhasade, Akash, et al.
Published: (2026)
by: Dhasade, Akash, et al.
Published: (2026)
One-Hop Sub-Query Result Caches for Graph Database Systems
by: Nguyen, Hieu, et al.
Published: (2024)
by: Nguyen, Hieu, et al.
Published: (2024)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
by: Bando, Yosuke, et al.
Published: (2025)
by: Bando, Yosuke, et al.
Published: (2025)
Sig2Model: A Boosting-Driven Model for Updatable Learned Indexes
by: Heidari, Alireza, et al.
Published: (2025)
by: Heidari, Alireza, et al.
Published: (2025)
VDTuner: Automated Performance Tuning for Vector Data Management Systems
by: Yang, Tiannuo, et al.
Published: (2024)
by: Yang, Tiannuo, et al.
Published: (2024)
Can Graph Reordering Speed Up Graph Neural Network Training? An Experimental Study
by: Merkel, Nikolai, et al.
Published: (2024)
by: Merkel, Nikolai, et al.
Published: (2024)
P-MOSS: Scheduling Main-Memory Indexes Over NUMA Servers Using Next Token Prediction
by: Rayhan, Yeasir, et al.
Published: (2024)
by: Rayhan, Yeasir, et al.
Published: (2024)
VectorSearch: Enhancing Document Retrieval with Semantic Embeddings and Optimized Search
by: Monir, Solmaz Seyed, et al.
Published: (2024)
by: Monir, Solmaz Seyed, et al.
Published: (2024)
On-Demand JSON: A Better Way to Parse Documents?
by: Keiser, John, et al.
Published: (2023)
by: Keiser, John, et al.
Published: (2023)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
by: Ulla, Aman
Published: (2026)
by: Ulla, Aman
Published: (2026)
Implementing and Evaluating E2LSH on Storage
by: Nakanishi, Yu, et al.
Published: (2024)
by: Nakanishi, Yu, et al.
Published: (2024)
On the effects of logical database design on database size, query complexity, query performance, and energy consumption
by: Taipalus, Toni
Published: (2025)
by: Taipalus, Toni
Published: (2025)
Robust Recursive Query Parallelism in Graph Database Management Systems
by: Chakraborty, Anurag, et al.
Published: (2025)
by: Chakraborty, Anurag, et al.
Published: (2025)
Micro-architectural Analysis of OLAP: Limitations and Opportunities
by: Sirin, Utku, et al.
Published: (2019)
by: Sirin, Utku, et al.
Published: (2019)
HistogramTools for Efficient Data Analysis and Distribution Representation in Large Data Sets
by: Malhotra, Shubham
Published: (2025)
by: Malhotra, Shubham
Published: (2025)
Accelerating Machine Learning Queries with Linear Algebra Query Processing
by: Sun, Wenbo, et al.
Published: (2023)
by: Sun, Wenbo, et al.
Published: (2023)
A Modular Graph-Native Query Optimization Framework
by: Lyu, Bingqing, et al.
Published: (2024)
by: Lyu, Bingqing, et al.
Published: (2024)
CXL and the Return of Scale-Up Database Engines
by: Lerner, Alberto, et al.
Published: (2024)
by: Lerner, Alberto, et al.
Published: (2024)
Parsing Gigabytes of JSON per Second
by: Langdale, Geoff, et al.
Published: (2019)
by: Langdale, Geoff, et al.
Published: (2019)
BranchBench: Aligning Database Branching with Agentic Demands
by: Ang, Elaine, et al.
Published: (2026)
by: Ang, Elaine, et al.
Published: (2026)
GraphMini: Accelerating Graph Pattern Matching Using Auxiliary Graphs
by: Liu, Juelin, et al.
Published: (2024)
by: Liu, Juelin, et al.
Published: (2024)
FLEXIS: FLEXible Frequent Subgraph Mining using Maximal Independent Sets
by: Sharma, Akshit, et al.
Published: (2024)
by: Sharma, Akshit, et al.
Published: (2024)
Heterogeneous Data Access Model for Concurrency Control and Methods to Deal with High Data Contention
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
by: Barad, Haim, et al.
Published: (2023)
by: Barad, Haim, et al.
Published: (2023)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
by: Tajdari, Sabiha, et al.
Published: (2025)
by: Tajdari, Sabiha, et al.
Published: (2025)
Real-time Event Joining in Practice With Kafka and Flink
by: Saket, Srijan, et al.
Published: (2024)
by: Saket, Srijan, et al.
Published: (2024)
RIVA: Leveraging LLM Agents for Reliable Configuration Drift Detection
by: Abuzakuk, Sami, et al.
Published: (2026)
by: Abuzakuk, Sami, et al.
Published: (2026)
Harnessing Increased Client Participation with Cohort-Parallel Federated Learning
by: Dhasade, Akash, et al.
Published: (2024)
by: Dhasade, Akash, et al.
Published: (2024)
Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
by: Jain, Rishabh, et al.
Published: (2024)
by: Jain, Rishabh, et al.
Published: (2024)
Spark Policy Toolkit: Semantic Contracts and Scalable Execution for Policy Learning in Spark
by: Bai, Zeyu
Published: (2026)
by: Bai, Zeyu
Published: (2026)
FB$^+$-tree: A Memory-Optimized B$^+$-tree with Latch-Free Update
by: Chen, Yuan, et al.
Published: (2025)
by: Chen, Yuan, et al.
Published: (2025)
CAMP: A Cost Adaptive Multi-Queue Eviction Policy for Key-Value Stores
by: Ghandeharizadeh, Shahram, et al.
Published: (2024)
by: Ghandeharizadeh, Shahram, et al.
Published: (2024)
AR-PPF: Advanced Resolution-Based Pixel Preemption Data Filtering for Efficient Time-Series Data Analysis
by: Kim, Taewoong, et al.
Published: (2024)
by: Kim, Taewoong, et al.
Published: (2024)
PHast -- Perfect Hashing made fast
by: Beling, Piotr, et al.
Published: (2025)
by: Beling, Piotr, et al.
Published: (2025)
Adaptive Hybrid Sort: Dynamic Strategy Selection for Optimal Sorting Across Diverse Data Distributions
by: Balasubramanian, Shrinivass Arunachalam
Published: (2025)
by: Balasubramanian, Shrinivass Arunachalam
Published: (2025)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Accelerating Transfer Learning with Near-Data Computation on Cloud Object Stores
by: Petrescu, Diana, et al.
Published: (2022)
by: Petrescu, Diana, et al.
Published: (2022)
Terabyte-Scale Analytics in the Blink of an Eye
by: Wu, Bowen, et al.
Published: (2025)
by: Wu, Bowen, et al.
Published: (2025)
Similar Items
-
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025) -
Catapults to the Rescue: Accelerating Vector Search by Exploiting Query Locality
by: Abuzakuk, Sami, et al.
Published: (2026) -
Effective LoRA Adapter Routing using Task Representations
by: Dhasade, Akash, et al.
Published: (2026) -
One-Hop Sub-Query Result Caches for Graph Database Systems
by: Nguyen, Hieu, et al.
Published: (2024) -
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
by: Bando, Yosuke, et al.
Published: (2025)