KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Huawei, Xia, Chunwei, Wang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
Leyline: KV Cache Directives for Agentic Inference
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model Acceleration based on KV Cache Management
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
von: Liu, Yanyu, et al.
Veröffentlicht: (2025)
von: Liu, Yanyu, et al.
Veröffentlicht: (2025)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
A Model Aware AIGC Task Offloading Algorithm in IIoT Edge Computing
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
Joint Resource Optimization, Computation Offloading and Resource Slicing for Multi-Edge Traffic-Cognitive Networks
von: Xiaoyang, Ting, et al.
Veröffentlicht: (2024)
von: Xiaoyang, Ting, et al.
Veröffentlicht: (2024)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
von: Moothedath, Vishnu Narayanan, et al.
Veröffentlicht: (2025)
von: Moothedath, Vishnu Narayanan, et al.
Veröffentlicht: (2025)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
von: Meng, William, et al.
Veröffentlicht: (2025)
von: Meng, William, et al.
Veröffentlicht: (2025)
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
von: Yildiz, Mert, et al.
Veröffentlicht: (2026)
von: Yildiz, Mert, et al.
Veröffentlicht: (2026)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
MatKV: Trading Compute for Flash Storage in LLM Inference
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
von: Li, Allison, et al.
Veröffentlicht: (2025)
von: Li, Allison, et al.
Veröffentlicht: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026) -
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025) -
Leyline: KV Cache Directives for Agentic Inference
von: Ma, Bole, et al.
Veröffentlicht: (2026) -
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
von: Li, Weiqing, et al.
Veröffentlicht: (2025) -
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025)