SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hongyao, Zhai, Liuqun, Wang, Junyi, Fang, Zhengru |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Efficient Wireless iBCI Headstage with Adaptive ADC Sample Rate
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
Offloading and Quality Control for AI Generated Content Services in 6G Mobile Edge Computing Networks
by: Wang, Yitong, et al.
Published: (2023)
by: Wang, Yitong, et al.
Published: (2023)
Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices
by: Yan, Xiao, et al.
Published: (2025)
by: Yan, Xiao, et al.
Published: (2025)
An Introductory Study on the Power Consumption Overhead of ERC-4337 Bundlers
by: Arusoaie, Andrei, et al.
Published: (2025)
by: Arusoaie, Andrei, et al.
Published: (2025)
Offline Reinforcement Learning for Mobility Robustness Optimization
by: Alizadeh, Pegah, et al.
Published: (2025)
by: Alizadeh, Pegah, et al.
Published: (2025)
SLA-Aware Distributed LLM Inference Across Device-RAN-Cloud
by: Yet, Hariz, et al.
Published: (2026)
by: Yet, Hariz, et al.
Published: (2026)
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
by: Moslem, Yasmin, et al.
Published: (2026)
by: Moslem, Yasmin, et al.
Published: (2026)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
by: Hoefler, Torsten, et al.
Published: (2026)
by: Hoefler, Torsten, et al.
Published: (2026)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
by: Jang, SiYoung, et al.
Published: (2025)
by: Jang, SiYoung, et al.
Published: (2025)
CarbonCP: Carbon-Aware DNN Partitioning with Conformal Prediction for Sustainable Edge Intelligence
by: Ke, Hongyu, et al.
Published: (2024)
by: Ke, Hongyu, et al.
Published: (2024)
Through the Lens of Google CrUX: Dissecting Web Browsing Experience Across Devices and Countries
by: Sengupta, Jayasree, et al.
Published: (2023)
by: Sengupta, Jayasree, et al.
Published: (2023)
LLM-Empowered Cooperative Content Caching in Vehicular Fog Caching-Assisted Platoon Networks
by: Tan, Bowen, et al.
Published: (2026)
by: Tan, Bowen, et al.
Published: (2026)
Unveiling Energy Efficiency in Deep Learning: Measurement, Prediction, and Scoring across Edge Devices
by: Tu, Xiaolong, et al.
Published: (2023)
by: Tu, Xiaolong, et al.
Published: (2023)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
PartialLoading: User Scheduling and Bandwidth Allocation for Parameter-sharing Edge Inference
by: Qu, Guanqiao, et al.
Published: (2025)
by: Qu, Guanqiao, et al.
Published: (2025)
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
by: Guan, Yongjie
Published: (2025)
by: Guan, Yongjie
Published: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
On Efficient Topology Management in Service-Oriented 6G Networks: An Edge Video Distribution Case Study
by: Ennaceur, Zied, et al.
Published: (2024)
by: Ennaceur, Zied, et al.
Published: (2024)
XLB: A High Performance Layer-7 Load Balancer for Microservices using eBPF-based In-kernel Interposition
by: Wang, Yuejie, et al.
Published: (2026)
by: Wang, Yuejie, et al.
Published: (2026)
Consistent Channel Hopping Algorithms for the Multichannel Rendezvous Problem with Heterogeneous Available Channel Sets
by: Liu, Yiwei, et al.
Published: (2025)
by: Liu, Yiwei, et al.
Published: (2025)
TrimCaching: Parameter-sharing AI Model Caching in Wireless Edge Networks
by: Qu, Guanqiao, et al.
Published: (2024)
by: Qu, Guanqiao, et al.
Published: (2024)
Analysis of Robust and Secure DNS Protocols for IoT Devices
by: Aydeger, Abdullah, et al.
Published: (2025)
by: Aydeger, Abdullah, et al.
Published: (2025)
Determinação Automática de Limiar de Detecção de Ataques em Redes de Computadores Utilizando Autoencoders
by: Miranda, Luan Gonçalves, et al.
Published: (2025)
by: Miranda, Luan Gonçalves, et al.
Published: (2025)
Optimizing LoRa for Edge Computing with TinyML Pipeline for Channel Hopping
by: Grunewald, Marla, et al.
Published: (2024)
by: Grunewald, Marla, et al.
Published: (2024)
Unleashing Automated Congestion Control Customization in the Wild
by: Cohen, Amit, et al.
Published: (2025)
by: Cohen, Amit, et al.
Published: (2025)
Selecting Offline Reinforcement Learning Algorithms for Stochastic Network Control
by: Helson, Nicolas, et al.
Published: (2026)
by: Helson, Nicolas, et al.
Published: (2026)
Generative AI on the Edge: Architecture and Performance Evaluation
by: Nezami, Zeinab, et al.
Published: (2024)
by: Nezami, Zeinab, et al.
Published: (2024)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
Enhancements to P4TG: Histogram-Based RTT Monitoring in the Data Plane
by: Ihle, Fabian, et al.
Published: (2025)
by: Ihle, Fabian, et al.
Published: (2025)
A protocol to reduce worst-case latency in deflection-based on-chip networks
by: Indrusiak, Leandro Soares
Published: (2025)
by: Indrusiak, Leandro Soares
Published: (2025)
Starlink on the Road: A First Look at Mobile Starlink Performance in Central Europe
by: Laniewski, Dominic, et al.
Published: (2024)
by: Laniewski, Dominic, et al.
Published: (2024)
NR Cell Identity-based Handover Decision-making Algorithm for High-speed Scenario within Dual Connectivity
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
Impact of Packetization on Network Calculus Analysis
by: Jiang, Yming
Published: (2025)
by: Jiang, Yming
Published: (2025)
Preprocess your Paths -- Speeding up Linear Programming-based Optimization for Segment Routing Traffic Engineering
by: Brundiers, Alexander, et al.
Published: (2023)
by: Brundiers, Alexander, et al.
Published: (2023)
An In-Depth Investigation of LEO Satellite Topology Design Parameters
by: Zhang, Wenyi, et al.
Published: (2024)
by: Zhang, Wenyi, et al.
Published: (2024)
Similar Items
-
An Efficient Wireless iBCI Headstage with Adaptive ADC Sample Rate
by: Liu, Hongyao, et al.
Published: (2026) -
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025) -
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026) -
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025) -
Offloading and Quality Control for AI Generated Content Services in 6G Mobile Edge Computing Networks
by: Wang, Yitong, et al.
Published: (2023)