Efficient Remote KV Cache Reuse with GPU-native Video Codec
Fuente:
arXiv
Salvato in:
| Autori principali: | Mi, Liang, Wang, Weijun, Chen, Jinghan, Cao, Ting, Dai, Haipeng, Liu, Yunxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
di: Li, Allison, et al.
Pubblicazione: (2025)
di: Li, Allison, et al.
Pubblicazione: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
di: Zhong, Shuzhang, et al.
Pubblicazione: (2025)
di: Zhong, Shuzhang, et al.
Pubblicazione: (2025)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
di: Lee, Wonbeom, et al.
Pubblicazione: (2024)
di: Lee, Wonbeom, et al.
Pubblicazione: (2024)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
di: Wang, Shao, et al.
Pubblicazione: (2026)
di: Wang, Shao, et al.
Pubblicazione: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
di: Qianli, Liu, et al.
Pubblicazione: (2025)
di: Qianli, Liu, et al.
Pubblicazione: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
di: Wang, Tianze, et al.
Pubblicazione: (2025)
di: Wang, Tianze, et al.
Pubblicazione: (2025)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
Leyline: KV Cache Directives for Agentic Inference
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
VcLLM: Video Codecs are Secretly Tensor Codecs
di: Xu, Ceyu, et al.
Pubblicazione: (2024)
di: Xu, Ceyu, et al.
Pubblicazione: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
di: Zou, Yulin, et al.
Pubblicazione: (2026)
di: Zou, Yulin, et al.
Pubblicazione: (2026)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
di: Zhang, Yanqi, et al.
Pubblicazione: (2024)
di: Zhang, Yanqi, et al.
Pubblicazione: (2024)
Decentralized Federated Learning with Model Caching on Mobile Agents
di: Wang, Xiaoyu, et al.
Pubblicazione: (2024)
di: Wang, Xiaoyu, et al.
Pubblicazione: (2024)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
di: Patel, Ishan, et al.
Pubblicazione: (2026)
di: Patel, Ishan, et al.
Pubblicazione: (2026)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
di: Li, Jinhao, et al.
Pubblicazione: (2023)
di: Li, Jinhao, et al.
Pubblicazione: (2023)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
di: Griggs, Tyler, et al.
Pubblicazione: (2024)
di: Griggs, Tyler, et al.
Pubblicazione: (2024)
Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
di: Peng, Jie, et al.
Pubblicazione: (2024)
di: Peng, Jie, et al.
Pubblicazione: (2024)
TDC-Cache: A Trustworthy Decentralized Cooperative Caching Framework for Web3.0
di: Chen, Jinyu, et al.
Pubblicazione: (2025)
di: Chen, Jinyu, et al.
Pubblicazione: (2025)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
di: Yu, Zhongming, et al.
Pubblicazione: (2024)
di: Yu, Zhongming, et al.
Pubblicazione: (2024)
GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching
di: Bathie, Gabriel, et al.
Pubblicazione: (2026)
di: Bathie, Gabriel, et al.
Pubblicazione: (2026)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
di: Nian, Sean, et al.
Pubblicazione: (2026)
di: Nian, Sean, et al.
Pubblicazione: (2026)
Keras Sig: Efficient Path Signature Computation on GPU in Keras 3
di: Genet, Rémi, et al.
Pubblicazione: (2025)
di: Genet, Rémi, et al.
Pubblicazione: (2025)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
di: Wu, Shixun, et al.
Pubblicazione: (2024)
di: Wu, Shixun, et al.
Pubblicazione: (2024)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
di: Zhao, Bingzhe, et al.
Pubblicazione: (2025)
di: Zhao, Bingzhe, et al.
Pubblicazione: (2025)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
di: Afroz, Sabiha, et al.
Pubblicazione: (2025)
di: Afroz, Sabiha, et al.
Pubblicazione: (2025)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
di: Trifan, Octavian Alexandru, et al.
Pubblicazione: (2025)
di: Trifan, Octavian Alexandru, et al.
Pubblicazione: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure
di: Stepanek, Lukas
Pubblicazione: (2026)
di: Stepanek, Lukas
Pubblicazione: (2026)
Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
di: Qin, Jianxing, et al.
Pubblicazione: (2025)
di: Qin, Jianxing, et al.
Pubblicazione: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
FedGroup: Efficient Clustered Federated Learning via Decomposed Data-Driven Measure
di: Duan, Moming, et al.
Pubblicazione: (2020)
di: Duan, Moming, et al.
Pubblicazione: (2020)
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
di: Jin, Chao, et al.
Pubblicazione: (2024)
di: Jin, Chao, et al.
Pubblicazione: (2024)
KV Cache Compression for Inference Efficiency in LLMs: A Review
di: Liu, Yanyu, et al.
Pubblicazione: (2025)
di: Liu, Yanyu, et al.
Pubblicazione: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
di: Luo, Ziyue, et al.
Pubblicazione: (2025)
di: Luo, Ziyue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
di: Li, Allison, et al.
Pubblicazione: (2025) -
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
di: Kim, Kihyun, et al.
Pubblicazione: (2025) -
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
di: Zhong, Shuzhang, et al.
Pubblicazione: (2025) -
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
di: Lee, Wonbeom, et al.
Pubblicazione: (2024) -
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
di: Wang, Shao, et al.
Pubblicazione: (2026)