LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xiong, Yi, Wu, Hao, Shao, Changxu, Wang, Ziqing, Zhang, Rui, Guo, Yuhong, Zhao, Junping, Zhang, Ke, Pan, Zhenxuan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
par: Bian, Zhuohang, et autres
Publié: (2025)
par: Bian, Zhuohang, et autres
Publié: (2025)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
par: Saadati, Yasaman, et autres
Publié: (2024)
par: Saadati, Yasaman, et autres
Publié: (2024)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
par: Wang, Shao, et autres
Publié: (2026)
par: Wang, Shao, et autres
Publié: (2026)
FedPLT: Scalable, Resource-Efficient, and Heterogeneity-Aware Federated Learning via Partial Layer Training
par: Dabaja, Ahmad, et autres
Publié: (2026)
par: Dabaja, Ahmad, et autres
Publié: (2026)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
par: Zhang, Haoran, et autres
Publié: (2025)
par: Zhang, Haoran, et autres
Publié: (2025)
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
par: Li, Zonghang, et autres
Publié: (2024)
par: Li, Zonghang, et autres
Publié: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
par: Su, Zhaoyuan, et autres
Publié: (2025)
par: Su, Zhaoyuan, et autres
Publié: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
par: Cox, Bart, et autres
Publié: (2022)
par: Cox, Bart, et autres
Publié: (2022)
Roadmap for Edge AI: A Dagstuhl Perspective
par: Ding, Aaron Yi, et autres
Publié: (2021)
par: Ding, Aaron Yi, et autres
Publié: (2021)
Parameterizing Federated Continual Learning for Reproducible Research
par: Cox, Bart, et autres
Publié: (2024)
par: Cox, Bart, et autres
Publié: (2024)
Asynchronous Byzantine Federated Learning
par: Cox, Bart, et autres
Publié: (2024)
par: Cox, Bart, et autres
Publié: (2024)
Training Diffusion Models with Federated Learning
par: de Goede, Matthijs, et autres
Publié: (2024)
par: de Goede, Matthijs, et autres
Publié: (2024)
Connecting Large Language Model Agent to High Performance Computing Resource
par: Ma, Heng, et autres
Publié: (2025)
par: Ma, Heng, et autres
Publié: (2025)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
par: Xin, Jihao, et autres
Publié: (2023)
par: Xin, Jihao, et autres
Publié: (2023)
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
par: Schroeder, Christian, et autres
Publié: (2023)
par: Schroeder, Christian, et autres
Publié: (2023)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
par: Zuo, Yuncong, et autres
Publié: (2024)
par: Zuo, Yuncong, et autres
Publié: (2024)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
par: Nian, Sean, et autres
Publié: (2026)
par: Nian, Sean, et autres
Publié: (2026)
SPARK: Igniting Communication-Efficient Decentralized Learning via Stage-wise Projected NTK and Accelerated Regularization
par: Xia, Li
Publié: (2025)
par: Xia, Li
Publié: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
par: Zhu, Jianian, et autres
Publié: (2025)
par: Zhu, Jianian, et autres
Publié: (2025)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
par: Wang, Zheng, et autres
Publié: (2025)
par: Wang, Zheng, et autres
Publié: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
par: Qianli, Liu, et autres
Publié: (2025)
par: Qianli, Liu, et autres
Publié: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
par: Bian, Zhuohang, et autres
Publié: (2026)
par: Bian, Zhuohang, et autres
Publié: (2026)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
par: Xu, Youxuan, et autres
Publié: (2025)
par: Xu, Youxuan, et autres
Publié: (2025)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
par: Chen, Jinfan, et autres
Publié: (2023)
par: Chen, Jinfan, et autres
Publié: (2023)
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
par: Zhao, Mingjie, et autres
Publié: (2026)
par: Zhao, Mingjie, et autres
Publié: (2026)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
par: Murimi, Almond Kiruthu
Publié: (2025)
par: Murimi, Almond Kiruthu
Publié: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
par: Zhang, Guilin, et autres
Publié: (2025)
par: Zhang, Guilin, et autres
Publié: (2025)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
par: Wang, Xiaoyu, et autres
Publié: (2025)
par: Wang, Xiaoyu, et autres
Publié: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
par: Petrov, Ivo, et autres
Publié: (2024)
par: Petrov, Ivo, et autres
Publié: (2024)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
par: Pugachev, Sergey
Publié: (2025)
par: Pugachev, Sergey
Publié: (2025)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
par: Parmar, Abhinav Singh
Publié: (2026)
par: Parmar, Abhinav Singh
Publié: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
par: Yoon, Dongha, et autres
Publié: (2025)
par: Yoon, Dongha, et autres
Publié: (2025)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
par: Zou, Jing, et autres
Publié: (2026)
par: Zou, Jing, et autres
Publié: (2026)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2024)
par: Xu, Jiale, et autres
Publié: (2024)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
par: Krishna, Gutha Jaya
Publié: (2024)
par: Krishna, Gutha Jaya
Publié: (2024)
Uncertainty Estimation in Multi-Agent Distributed Learning for AI-Enabled Edge Devices
par: Radchenko, Gleb, et autres
Publié: (2024)
par: Radchenko, Gleb, et autres
Publié: (2024)
Edge AI Collaborative Learning: Bayesian Approaches to Uncertainty Estimation
par: Radchenko, Gleb, et autres
Publié: (2024)
par: Radchenko, Gleb, et autres
Publié: (2024)
SPEAR++: Scaling Gradient Inversion via Sparsely-Used Dictionary Learning
par: Bakarsky, Alexander, et autres
Publié: (2025)
par: Bakarsky, Alexander, et autres
Publié: (2025)
Documents similaires
-
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
par: Bian, Zhuohang, et autres
Publié: (2025) -
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
par: Saadati, Yasaman, et autres
Publié: (2024) -
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
par: Wang, Shao, et autres
Publié: (2026) -
FedPLT: Scalable, Resource-Efficient, and Heterogeneity-Aware Federated Learning via Partial Layer Training
par: Dabaja, Ahmad, et autres
Publié: (2026) -
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
par: Zhang, Haoran, et autres
Publié: (2025)