DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Strati, Foteini, Mcallister, Sara, Phanishayee, Amar, Tarnawski, Jakub, Klimovic, Ana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DejaVu: A Minimalistic Mechanism for Distributed Plurality Consensus
di: d'Amore, Francesco, et al.
Pubblicazione: (2026)
di: d'Amore, Francesco, et al.
Pubblicazione: (2026)
Understanding GPU Resource Interference One Level Deeper
di: Elvinger, Paul, et al.
Pubblicazione: (2025)
di: Elvinger, Paul, et al.
Pubblicazione: (2025)
Integrated Hardware Architecture and Device Placement Search
di: Wang, Irene, et al.
Pubblicazione: (2024)
di: Wang, Irene, et al.
Pubblicazione: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
di: Strati, Foteini, et al.
Pubblicazione: (2025)
di: Strati, Foteini, et al.
Pubblicazione: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
di: Yao, Xiaozhe, et al.
Pubblicazione: (2023)
di: Yao, Xiaozhe, et al.
Pubblicazione: (2023)
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
di: Giannoula, Christina, et al.
Pubblicazione: (2024)
di: Giannoula, Christina, et al.
Pubblicazione: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
di: Stepanek, Lukas
Pubblicazione: (2026)
di: Stepanek, Lukas
Pubblicazione: (2026)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
di: Nian, Sean, et al.
Pubblicazione: (2026)
di: Nian, Sean, et al.
Pubblicazione: (2026)
Fast State Restoration in LLM Serving with HCache
di: Gao, Shiwei, et al.
Pubblicazione: (2024)
di: Gao, Shiwei, et al.
Pubblicazione: (2024)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
di: Hwang, Jinwoo, et al.
Pubblicazione: (2025)
di: Hwang, Jinwoo, et al.
Pubblicazione: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
di: Adnan, Muhammad, et al.
Pubblicazione: (2024)
di: Adnan, Muhammad, et al.
Pubblicazione: (2024)
Fast and Interactive Byzantine Fault-tolerant Web Services via Session-Based Consensus Decoupling
di: Akmal, Ahmad Zaki, et al.
Pubblicazione: (2025)
di: Akmal, Ahmad Zaki, et al.
Pubblicazione: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
di: Guo, Yipin, et al.
Pubblicazione: (2026)
di: Guo, Yipin, et al.
Pubblicazione: (2026)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
Fault-tolerant Consensus in Anonymous Dynamic Network
di: Zhang, Qinzi, et al.
Pubblicazione: (2024)
di: Zhang, Qinzi, et al.
Pubblicazione: (2024)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
Fault-tolerant Reduce and Allreduce operations based on correction
di: Kuettler, Martin, et al.
Pubblicazione: (2026)
di: Kuettler, Martin, et al.
Pubblicazione: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
Hamava: Fault-tolerant Reconfigurable Geo-Replication on Heterogeneous Clusters
di: Mane, Tejas, et al.
Pubblicazione: (2024)
di: Mane, Tejas, et al.
Pubblicazione: (2024)
The Power of Abstract MAC Layer: A Fault-tolerance Perspective
di: Zhang, Qinzi, et al.
Pubblicazione: (2024)
di: Zhang, Qinzi, et al.
Pubblicazione: (2024)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
di: Liu, Lian, et al.
Pubblicazione: (2026)
di: Liu, Lian, et al.
Pubblicazione: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
di: Qianli, Liu, et al.
Pubblicazione: (2025)
di: Qianli, Liu, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
di: Shen, Haiying, et al.
Pubblicazione: (2024)
di: Shen, Haiying, et al.
Pubblicazione: (2024)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
Predictable LLM Serving on GPU Clusters
di: Darzi, Erfan, et al.
Pubblicazione: (2025)
di: Darzi, Erfan, et al.
Pubblicazione: (2025)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DejaVu: A Minimalistic Mechanism for Distributed Plurality Consensus
di: d'Amore, Francesco, et al.
Pubblicazione: (2026) -
Understanding GPU Resource Interference One Level Deeper
di: Elvinger, Paul, et al.
Pubblicazione: (2025) -
Integrated Hardware Architecture and Device Placement Search
di: Wang, Irene, et al.
Pubblicazione: (2024) -
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
di: Jiang, Youhe, et al.
Pubblicazione: (2025) -
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
di: Strati, Foteini, et al.
Pubblicazione: (2025)