Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agarwal, Shubham, Sundaresan, Sai, Mitra, Subrata, Mahapatra, Debabrata, Gupta, Archit, Sharma, Rounak, Kapu, Nirmal Joshua, Yu, Tong, Saini, Shiv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025)
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025)
Demo-Craft: Using In-Context Learning to Improve Code Generation in Large Language Models
von: Kapu, Nirmal Joshua, et al.
Veröffentlicht: (2024)
von: Kapu, Nirmal Joshua, et al.
Veröffentlicht: (2024)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
Idiosyncrasies of Programmable Caching Engines
von: Peixoto, José, et al.
Veröffentlicht: (2026)
von: Peixoto, José, et al.
Veröffentlicht: (2026)
DPC: A Distributed Page Cache over CXL
von: Bergman, Shai, et al.
Veröffentlicht: (2026)
von: Bergman, Shai, et al.
Veröffentlicht: (2026)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
von: Berend, Daniel, et al.
Veröffentlicht: (2025)
von: Berend, Daniel, et al.
Veröffentlicht: (2025)
Random Adaptive Cache Placement Policy
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
von: Zhang, Dingyan, et al.
Veröffentlicht: (2024)
von: Zhang, Dingyan, et al.
Veröffentlicht: (2024)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
von: Zou, Jing, et al.
Veröffentlicht: (2026)
von: Zou, Jing, et al.
Veröffentlicht: (2026)
Revisiting Cache Freshness for Emerging Real-Time Applications
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
von: Kao, Henry, et al.
Veröffentlicht: (2025)
von: Kao, Henry, et al.
Veröffentlicht: (2025)
ONCache: A Cache-Based Low-Overhead Container Overlay Network
von: Lin, Shengkai, et al.
Veröffentlicht: (2023)
von: Lin, Shengkai, et al.
Veröffentlicht: (2023)
Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches
von: Cheng, Chiyu, et al.
Veröffentlicht: (2024)
von: Cheng, Chiyu, et al.
Veröffentlicht: (2024)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
Towards Advanced Speech Signal Processing: A Statistical Perspective on Convolution-Based Architectures and its Applications
von: Kapu, Nirmal Joshua, et al.
Veröffentlicht: (2024)
von: Kapu, Nirmal Joshua, et al.
Veröffentlicht: (2024)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
von: Lu, Songshuo, et al.
Veröffentlicht: (2024)
von: Lu, Songshuo, et al.
Veröffentlicht: (2024)
Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
von: Sun, Binqi, et al.
Veröffentlicht: (2025)
von: Sun, Binqi, et al.
Veröffentlicht: (2025)
Cache is King: Smart Page Eviction with eBPF
von: Zussman, Tal, et al.
Veröffentlicht: (2025)
von: Zussman, Tal, et al.
Veröffentlicht: (2025)
HACache: Leveraging Read Performance with Cache in a Heterogeneous Array
von: Liu, Jialin, et al.
Veröffentlicht: (2026)
von: Liu, Jialin, et al.
Veröffentlicht: (2026)
Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
2DIO: A Cache-Accurate Storage Microbenchmark
von: Wang, Yirong, et al.
Veröffentlicht: (2026)
von: Wang, Yirong, et al.
Veröffentlicht: (2026)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
von: Wu, Fangnuo, et al.
Veröffentlicht: (2025)
von: Wu, Fangnuo, et al.
Veröffentlicht: (2025)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
von: Qiu, Shi, et al.
Veröffentlicht: (2026)
von: Qiu, Shi, et al.
Veröffentlicht: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Efficient Text-to-Image Inferencing System
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2025)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
von: Nouri, Azam
Veröffentlicht: (2026)
von: Nouri, Azam
Veröffentlicht: (2026)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
LearnedCache: An eBPF-Integrated Perceptron-Based Eviction Policy for the Linux Page Cache
von: Qi, Zejia
Veröffentlicht: (2026)
von: Qi, Zejia
Veröffentlicht: (2026)
I/O Transit Caching for PMem-based Block Device
von: Xu, Qing, et al.
Veröffentlicht: (2024)
von: Xu, Qing, et al.
Veröffentlicht: (2024)
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
von: Lee, Kun-Hui, et al.
Veröffentlicht: (2025)
von: Lee, Kun-Hui, et al.
Veröffentlicht: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
von: Zeng, Yixiao, et al.
Veröffentlicht: (2026)
von: Zeng, Yixiao, et al.
Veröffentlicht: (2026)
Taming Serverless Cold Starts Through OS Co-Design
von: Holmes, Ben, et al.
Veröffentlicht: (2025)
von: Holmes, Ben, et al.
Veröffentlicht: (2025)
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
von: Jin, Chao, et al.
Veröffentlicht: (2024)
von: Jin, Chao, et al.
Veröffentlicht: (2024)
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
von: Kanellopoulos, Konstantinos, et al.
Veröffentlicht: (2023)
von: Kanellopoulos, Konstantinos, et al.
Veröffentlicht: (2023)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025) -
Demo-Craft: Using In-Context Learning to Improve Code Generation in Large Language Models
von: Kapu, Nirmal Joshua, et al.
Veröffentlicht: (2024) -
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
von: Tofigh, Mani, et al.
Veröffentlicht: (2025) -
Idiosyncrasies of Programmable Caching Engines
von: Peixoto, José, et al.
Veröffentlicht: (2026) -
DPC: A Distributed Page Cache over CXL
von: Bergman, Shai, et al.
Veröffentlicht: (2026)