Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hanchen, Liu, Yuhan, Cheng, Yihua, Du, Kuntai, Jiang, Junchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
by: Li, Hanchen, et al.
Published: (2024)
by: Li, Hanchen, et al.
Published: (2024)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
Earth+: on-board satellite imagery compression leveraging historical earth observations
by: Du, Kuntai, et al.
Published: (2024)
by: Du, Kuntai, et al.
Published: (2024)
Loss-tolerant neural video codec aware congestion control for real time video communication
by: Xia, Zhengxu, et al.
Published: (2024)
by: Xia, Zhengxu, et al.
Published: (2024)
Pushing the Limits of In-Network Caching for Key-Value Stores
by: Kim, Gyuyeong
Published: (2024)
by: Kim, Gyuyeong
Published: (2024)
VNF-Cache: An In-Network Key-Value Store Cache Based on Network Function Virtualization
by: Farias, Bruno E., et al.
Published: (2025)
by: Farias, Bruno E., et al.
Published: (2025)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
5GC$^2$ache: Improving 5G UPF Performance via Cache Optimization
by: Jia, Haonan, et al.
Published: (2024)
by: Jia, Haonan, et al.
Published: (2024)
GRACE: Loss-Resilient Real-Time Video through Neural Codecs
by: Cheng, Yihua, et al.
Published: (2023)
by: Cheng, Yihua, et al.
Published: (2023)
FluxShard: Motion-Aware Feature Cache Reuse for Collaborative Video Analytics in Mobile Edge Computing
by: Guan, Xiuxian, et al.
Published: (2026)
by: Guan, Xiuxian, et al.
Published: (2026)
Model Context Protocol-based Internet of Experts For Wireless Environment-aware LLM Agents
by: Liu, Zongxi, et al.
Published: (2025)
by: Liu, Zongxi, et al.
Published: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
NetMCP: Network-Aware Model Context Protocol Platform for LLM Capability Extension
by: Li, Enhan, et al.
Published: (2025)
by: Li, Enhan, et al.
Published: (2025)
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
by: Du, Kuntai, et al.
Published: (2023)
by: Du, Kuntai, et al.
Published: (2023)
Online Digital Twin-Empowered Content Resale Mechanism in Age of Information-Aware Edge Caching Networks
by: Yi, Yuhan, et al.
Published: (2024)
by: Yi, Yuhan, et al.
Published: (2024)
TrimCaching: Parameter-sharing Edge Caching for AI Model Downloading
by: Qu, Guanqiao, et al.
Published: (2024)
by: Qu, Guanqiao, et al.
Published: (2024)
Adaptive Contextual Caching for Mobile Edge Large Language Model Service
by: Liu, Guangyuan, et al.
Published: (2025)
by: Liu, Guangyuan, et al.
Published: (2025)
M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
by: Zeng, Yongjie, et al.
Published: (2025)
by: Zeng, Yongjie, et al.
Published: (2025)
LLM-Empowered Cooperative Content Caching in Vehicular Fog Caching-Assisted Platoon Networks
by: Tan, Bowen, et al.
Published: (2026)
by: Tan, Bowen, et al.
Published: (2026)
Modeling and Optimizing Latency for Delayed Hit Caching with Stochastic Miss Latency
by: Jiang, Bowen, et al.
Published: (2025)
by: Jiang, Bowen, et al.
Published: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
Semantic-Aware Caching for Efficient Image Generation in Edge Computing
by: Cui, Hanshuai, et al.
Published: (2025)
by: Cui, Hanshuai, et al.
Published: (2025)
Learning Cache Coherence Traffic for NoC Routing Design
by: Xiong, Guochu, et al.
Published: (2025)
by: Xiong, Guochu, et al.
Published: (2025)
Joint Model Caching and Resource Allocation in Generative AI-Enabled Wireless Edge Networks
by: Liu, Zhang, et al.
Published: (2024)
by: Liu, Zhang, et al.
Published: (2024)
Serving Long-Context LLMs at the Mobile Edge: Test-Time Reinforcement Learning-based Model Caching and Inference Offloading
by: Xu, Minrui, et al.
Published: (2025)
by: Xu, Minrui, et al.
Published: (2025)
VA-CDH: A Variance-Aware Method to Optimize Latency for Caching with Delayed Hits
by: Jiang, Bowen, et al.
Published: (2025)
by: Jiang, Bowen, et al.
Published: (2025)
Toward Generative 6G Simulation: An Experimental Multi-Agent LLM and ns-3 Integration
by: Rezazadeh, Farhad, et al.
Published: (2025)
by: Rezazadeh, Farhad, et al.
Published: (2025)
Improving Wi-Fi 8 Latency with Coordinated Spatial Reuse
by: Nunez, David, et al.
Published: (2025)
by: Nunez, David, et al.
Published: (2025)
NetLLM: Adapting Large Language Models for Networking
by: Wu, Duo, et al.
Published: (2024)
by: Wu, Duo, et al.
Published: (2024)
JAUNT: Joint Alignment of User Intent and Network State for QoE-centric LLM Tool Routing
by: Li, Enhan, et al.
Published: (2025)
by: Li, Enhan, et al.
Published: (2025)
CReIS: Computation Reuse through Image Similarity in ICN-Based Edge Computing
by: Javaheri, Atiyeh, et al.
Published: (2025)
by: Javaheri, Atiyeh, et al.
Published: (2025)
Fundamentals of Caching Layered Data objects
by: Bari, Agrim, et al.
Published: (2025)
by: Bari, Agrim, et al.
Published: (2025)
Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception
by: Liu, Guangyuan, et al.
Published: (2025)
by: Liu, Guangyuan, et al.
Published: (2025)
BLADE: Adaptive Wi-Fi Contention Control for Next-Generation Real-Time Communication
by: Guo, Fengqian, et al.
Published: (2026)
by: Guo, Fengqian, et al.
Published: (2026)
AgentVNE: LLM-Augmented Graph Reinforcement Learning for Affinity-Aware Multi-Agent Placement in Edge Agentic AI
by: Zheng, Runze, et al.
Published: (2026)
by: Zheng, Runze, et al.
Published: (2026)
Value-based Proactive Caching for Sensing Data in Vehicular Networks: An Operator's Perspective
by: Wang, Yantong, et al.
Published: (2024)
by: Wang, Yantong, et al.
Published: (2024)
Coordinated Spatial Reuse Scheduling With Machine Learning in IEEE 802.11 MAPC Networks
by: Wojnar, Maksymilian, et al.
Published: (2025)
by: Wojnar, Maksymilian, et al.
Published: (2025)
Digital Twin-Enabled Mobility-Aware Cooperative Caching in Vehicular Edge Computing
by: Zeng, Jiahao, et al.
Published: (2026)
by: Zeng, Jiahao, et al.
Published: (2026)
Similar Items
-
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
by: Li, Hanchen, et al.
Published: (2024) -
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023) -
Earth+: on-board satellite imagery compression leveraging historical earth observations
by: Du, Kuntai, et al.
Published: (2024) -
Loss-tolerant neural video codec aware congestion control for real time video communication
by: Xia, Zhengxu, et al.
Published: (2024) -
Pushing the Limits of In-Network Caching for Key-Value Stores
by: Kim, Gyuyeong
Published: (2024)