Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Allison, Greenewald, Kristjan, Parnell, Thomas, Azizan, Navid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
von: Wang, Shao, et al.
Veröffentlicht: (2026)
von: Wang, Shao, et al.
Veröffentlicht: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
von: Nian, Sean, et al.
Veröffentlicht: (2026)
von: Nian, Sean, et al.
Veröffentlicht: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
von: Bian, Zhuohang, et al.
Veröffentlicht: (2026)
von: Bian, Zhuohang, et al.
Veröffentlicht: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
FDLoRA: Personalized Federated Learning of Large Language Model via Dual LoRA Tuning
von: QI, Jiaxing, et al.
Veröffentlicht: (2024)
von: QI, Jiaxing, et al.
Veröffentlicht: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
von: Zhu, Jianian, et al.
Veröffentlicht: (2025)
von: Zhu, Jianian, et al.
Veröffentlicht: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
von: Mi, Liang, et al.
Veröffentlicht: (2026)
von: Mi, Liang, et al.
Veröffentlicht: (2026)
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
von: Liu, Jianmin, et al.
Veröffentlicht: (2025)
von: Liu, Jianmin, et al.
Veröffentlicht: (2025)
FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
von: Li, Rukuo, et al.
Veröffentlicht: (2025)
von: Li, Rukuo, et al.
Veröffentlicht: (2025)
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
von: Ding, Chuntao, et al.
Veröffentlicht: (2024)
von: Ding, Chuntao, et al.
Veröffentlicht: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
Federated LoRA with Sparse Communication
von: Kuo, Kevin, et al.
Veröffentlicht: (2024)
von: Kuo, Kevin, et al.
Veröffentlicht: (2024)
pFedLoRA: Model-Heterogeneous Personalized Federated Learning with LoRA Tuning
von: Yi, Liping, et al.
Veröffentlicht: (2023)
von: Yi, Liping, et al.
Veröffentlicht: (2023)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
von: Toniolo, Alessio Ricci, et al.
Veröffentlicht: (2026)
von: Toniolo, Alessio Ricci, et al.
Veröffentlicht: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024)
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
von: Ni, Yinan, et al.
Veröffentlicht: (2025)
von: Ni, Yinan, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
von: Meng, William, et al.
Veröffentlicht: (2025)
von: Meng, William, et al.
Veröffentlicht: (2025)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
von: Zuo, Jingwei, et al.
Veröffentlicht: (2026)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2026)
Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?
von: Wang, Yatong, et al.
Veröffentlicht: (2026)
von: Wang, Yatong, et al.
Veröffentlicht: (2026)
Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation Models
von: Cho, Yae Jee, et al.
Veröffentlicht: (2024)
von: Cho, Yae Jee, et al.
Veröffentlicht: (2024)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
von: Stepanek, Lukas
Veröffentlicht: (2026)
von: Stepanek, Lukas
Veröffentlicht: (2026)
Ähnliche Einträge
-
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024) -
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
von: Wang, Shao, et al.
Veröffentlicht: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023) -
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026) -
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)