CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Dong, Yu, Yanxuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
di: Shukla, Shikhar
Pubblicazione: (2026)
di: Shukla, Shikhar
Pubblicazione: (2026)
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
di: Liu, Zedong, et al.
Pubblicazione: (2026)
di: Liu, Zedong, et al.
Pubblicazione: (2026)
Joint Encoding of KV-Cache Blocks for Scalable LLM Serving
di: Kampeas, Joseph, et al.
Pubblicazione: (2026)
di: Kampeas, Joseph, et al.
Pubblicazione: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
di: Guo, Yipin, et al.
Pubblicazione: (2026)
di: Guo, Yipin, et al.
Pubblicazione: (2026)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
di: Li, Weiqing, et al.
Pubblicazione: (2025)
di: Li, Weiqing, et al.
Pubblicazione: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
di: Kim, Dowon, et al.
Pubblicazione: (2025)
di: Kim, Dowon, et al.
Pubblicazione: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2026)
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
di: Mamo, Oteo, et al.
Pubblicazione: (2026)
di: Mamo, Oteo, et al.
Pubblicazione: (2026)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
di: Jiang, Bo, et al.
Pubblicazione: (2025)
di: Jiang, Bo, et al.
Pubblicazione: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2024)
di: Feng, Yuan, et al.
Pubblicazione: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
di: Xiong, Yi, et al.
Pubblicazione: (2024)
di: Xiong, Yi, et al.
Pubblicazione: (2024)
The Pitfalls of KV Cache Compression
di: Chen, Alex, et al.
Pubblicazione: (2025)
di: Chen, Alex, et al.
Pubblicazione: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
di: Dong, Yanhao, et al.
Pubblicazione: (2025)
di: Dong, Yanhao, et al.
Pubblicazione: (2025)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
di: Li, Yubo, et al.
Pubblicazione: (2026)
di: Li, Yubo, et al.
Pubblicazione: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
di: Hao, Jitai, et al.
Pubblicazione: (2026)
di: Hao, Jitai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
di: Shukla, Shikhar
Pubblicazione: (2026) -
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025) -
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025) -
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
di: Liu, Zedong, et al.
Pubblicazione: (2026)