FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Dongwei, Liu, Zijie, Wang, Song, Ren, Yuxin, Deng, Jianing, Hu, Jingtong, Chen, Tianlong, Yang, Huanrui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
by: Deng, Jianing, et al.
Published: (2026)
by: Deng, Jianing, et al.
Published: (2026)
Dynamic Mixed-Precision Routing for Efficient Multi-step LLM Interaction
by: Li, Yuanzhe, et al.
Published: (2026)
by: Li, Yuanzhe, et al.
Published: (2026)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
by: Qi, Yanlin, et al.
Published: (2026)
by: Qi, Yanlin, et al.
Published: (2026)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
On 10x Better Scalability: KV Stores Scale Up KV Cache
by: Yu, Weiping, et al.
Published: (2025)
by: Yu, Weiping, et al.
Published: (2025)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
by: Yan, Jianxin, et al.
Published: (2026)
by: Yan, Jianxin, et al.
Published: (2026)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
by: Kim, Jang-Hyun, et al.
Published: (2025)
by: Kim, Jang-Hyun, et al.
Published: (2025)
gMatch: Fine-Grained and Hardware-Efficient Subgraph Matching on GPUs
by: Chen, Weitian, et al.
Published: (2026)
by: Chen, Weitian, et al.
Published: (2026)
Adaptive KV-Cache Compression without Manually Setting Budget
by: Tang, Chenxia, et al.
Published: (2025)
by: Tang, Chenxia, et al.
Published: (2025)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
by: Wu, Guanlong, et al.
Published: (2026)
by: Wu, Guanlong, et al.
Published: (2026)
AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
by: Deng, Yangshen, et al.
Published: (2025)
by: Deng, Yangshen, et al.
Published: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
Efficient Distributed Exact Subgraph Matching via GNN-PE: Load Balancing, Cache Optimization, and Query Plan Ranking
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
SynchroStore: A Cost-Based Fine-Grained Incremental Compaction for Hybrid Workloads
by: Zhang, Yinan, et al.
Published: (2025)
by: Zhang, Yinan, et al.
Published: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Vortex: Hosting ML Inference and Knowledge Retrieval Services With Tight Latency and Throughput Requirements
by: Yang, Yuting, et al.
Published: (2025)
by: Yang, Yuting, et al.
Published: (2025)
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025)
by: Jaillet, Patrick, et al.
Published: (2025)
GoVector: An I/O-Efficient Caching Strategy for High-Dimensional Vector Nearest Neighbor Search
by: Zhou, Yijie, et al.
Published: (2025)
by: Zhou, Yijie, et al.
Published: (2025)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Fine-Grained Table Retrieval Through the Lens of Complex Queries
by: Kosiuk, Wojciech, et al.
Published: (2026)
by: Kosiuk, Wojciech, et al.
Published: (2026)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
by: Bergman, Shai, et al.
Published: (2025)
by: Bergman, Shai, et al.
Published: (2025)
Taking the Leap: Efficient and Reliable Fine-Grained NUMA Migration in User-space
by: Schuhknecht, Felix, et al.
Published: (2026)
by: Schuhknecht, Felix, et al.
Published: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
by: Wu, Wenbo, et al.
Published: (2025)
by: Wu, Wenbo, et al.
Published: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025)
by: Tian, Yuxuan, et al.
Published: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
by: Shi, Zhiyuan, et al.
Published: (2026)
by: Shi, Zhiyuan, et al.
Published: (2026)
RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
by: Wang, Zhengchao, et al.
Published: (2025)
by: Wang, Zhengchao, et al.
Published: (2025)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
by: Yan, Jianxin, et al.
Published: (2025)
by: Yan, Jianxin, et al.
Published: (2025)
Compression and In-Situ Query Processing for Fine-Grained Array Lineage
by: Zhao, Jinjin, et al.
Published: (2024)
by: Zhao, Jinjin, et al.
Published: (2024)
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
by: Hu, Zhengding, et al.
Published: (2025)
by: Hu, Zhengding, et al.
Published: (2025)
HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL
by: Peng, You, et al.
Published: (2025)
by: Peng, You, et al.
Published: (2025)
RAC: Relation-Aware Cache Replacement for Large Language Models
by: Wu, Yuchong, et al.
Published: (2026)
by: Wu, Yuchong, et al.
Published: (2026)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
ZipCache: A DRAM/SSD Cache with Built-in Transparent Compression
by: Xie, Rui, et al.
Published: (2024)
by: Xie, Rui, et al.
Published: (2024)
BubbleRAG: Evidence-Driven Retrieval-Augmented Generation for Black-Box Knowledge Graphs
by: Pan, Duyi, et al.
Published: (2026)
by: Pan, Duyi, et al.
Published: (2026)
A Survey of LLM Inference Systems
by: Pan, James, et al.
Published: (2025)
by: Pan, James, et al.
Published: (2025)
Dynamic read & write optimization with TurtleKV
by: Astolfi, Tony, et al.
Published: (2025)
by: Astolfi, Tony, et al.
Published: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
CubeGraph: Efficient Retrieval-Augmented Generation for Spatial and Temporal Data
by: Yang, Mingyu, et al.
Published: (2026)
by: Yang, Mingyu, et al.
Published: (2026)
Similar Items
-
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
by: Deng, Jianing, et al.
Published: (2026) -
Dynamic Mixed-Precision Routing for Efficient Multi-step LLM Interaction
by: Li, Yuanzhe, et al.
Published: (2026) -
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
by: Qi, Yanlin, et al.
Published: (2026) -
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024) -
On 10x Better Scalability: KV Stores Scale Up KV Cache
by: Yu, Weiping, et al.
Published: (2025)