SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Quqing, Chen, Kai, Liao, Ning, Lin, Zehao, Tang, Bo, Xiong, Feiyu, Wang, Xiaoxing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
di: Lin, Yujun, et al.
Pubblicazione: (2024)
di: Lin, Yujun, et al.
Pubblicazione: (2024)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
di: Ma, Xinyue, et al.
Pubblicazione: (2026)
di: Ma, Xinyue, et al.
Pubblicazione: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
di: Zhang, Hang, et al.
Pubblicazione: (2025)
di: Zhang, Hang, et al.
Pubblicazione: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
di: Yang, Shang, et al.
Pubblicazione: (2025)
di: Yang, Shang, et al.
Pubblicazione: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
di: Yao, Feiyu, et al.
Pubblicazione: (2026)
di: Yao, Feiyu, et al.
Pubblicazione: (2026)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
di: Yu, Shan, et al.
Pubblicazione: (2025)
di: Yu, Shan, et al.
Pubblicazione: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
di: Liu, Minghui, et al.
Pubblicazione: (2025)
di: Liu, Minghui, et al.
Pubblicazione: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
di: Liu, Zirui, et al.
Pubblicazione: (2024)
di: Liu, Zirui, et al.
Pubblicazione: (2024)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
di: Hu, Junhao, et al.
Pubblicazione: (2024)
di: Hu, Junhao, et al.
Pubblicazione: (2024)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
di: Zhao, Youpeng, et al.
Pubblicazione: (2024)
di: Zhao, Youpeng, et al.
Pubblicazione: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
di: Barad, Haim, et al.
Pubblicazione: (2023)
di: Barad, Haim, et al.
Pubblicazione: (2023)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
Scaler: Efficient and Effective Cross Flow Analysis
di: Steven, et al.
Pubblicazione: (2024)
di: Steven, et al.
Pubblicazione: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
di: Hendria, Willy Fitra
Pubblicazione: (2026)
di: Hendria, Willy Fitra
Pubblicazione: (2026)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
di: Liu, Xiaoxuan, et al.
Pubblicazione: (2024)
di: Liu, Xiaoxuan, et al.
Pubblicazione: (2024)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
di: Du, Dayou, et al.
Pubblicazione: (2025)
di: Du, Dayou, et al.
Pubblicazione: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
di: Shin, Jiho, et al.
Pubblicazione: (2024)
di: Shin, Jiho, et al.
Pubblicazione: (2024)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies
di: Kim, Kyoungmin, et al.
Pubblicazione: (2024)
di: Kim, Kyoungmin, et al.
Pubblicazione: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
di: He, Jiaao, et al.
Pubblicazione: (2024)
di: He, Jiaao, et al.
Pubblicazione: (2024)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
di: Suo, Jiashun, et al.
Pubblicazione: (2025)
di: Suo, Jiashun, et al.
Pubblicazione: (2025)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
di: Yang, Chongzhuo, et al.
Pubblicazione: (2024)
di: Yang, Chongzhuo, et al.
Pubblicazione: (2024)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
di: Tu, Dezhan, et al.
Pubblicazione: (2024)
di: Tu, Dezhan, et al.
Pubblicazione: (2024)
Can Increasing the Hit Ratio Hurt Cache Throughput? (Long Version)
di: Qiu, Ziyue, et al.
Pubblicazione: (2024)
di: Qiu, Ziyue, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026) -
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024) -
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
di: Lin, Yujun, et al.
Pubblicazione: (2024) -
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
di: Ma, Xinyue, et al.
Pubblicazione: (2026) -
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
di: Zhang, Hang, et al.
Pubblicazione: (2025)