Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Allison, Greenewald, Kristjan, Parnell, Thomas, Azizan, Navid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024)
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
por: Wang, Shao, et al.
Publicado: (2026)
por: Wang, Shao, et al.
Publicado: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
por: Sheng, Ying, et al.
Publicado: (2023)
por: Sheng, Ying, et al.
Publicado: (2023)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
por: Chen, Hongyu, et al.
Publicado: (2026)
por: Chen, Hongyu, et al.
Publicado: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
por: Li, Suyi, et al.
Publicado: (2024)
por: Li, Suyi, et al.
Publicado: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
por: li, Fei, et al.
Publicado: (2026)
por: li, Fei, et al.
Publicado: (2026)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
por: Nian, Sean, et al.
Publicado: (2026)
por: Nian, Sean, et al.
Publicado: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
por: Qianli, Liu, et al.
Publicado: (2025)
por: Qianli, Liu, et al.
Publicado: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
por: Bian, Zhuohang, et al.
Publicado: (2026)
por: Bian, Zhuohang, et al.
Publicado: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
por: He, Yiyuan, et al.
Publicado: (2025)
por: He, Yiyuan, et al.
Publicado: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
por: Kim, Kihyun, et al.
Publicado: (2025)
por: Kim, Kihyun, et al.
Publicado: (2025)
FDLoRA: Personalized Federated Learning of Large Language Model via Dual LoRA Tuning
por: QI, Jiaxing, et al.
Publicado: (2024)
por: QI, Jiaxing, et al.
Publicado: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
por: Su, Zhaoyuan, et al.
Publicado: (2025)
por: Su, Zhaoyuan, et al.
Publicado: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
por: Zhu, Jianian, et al.
Publicado: (2025)
por: Zhu, Jianian, et al.
Publicado: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
por: Yoon, Dongha, et al.
Publicado: (2025)
por: Yoon, Dongha, et al.
Publicado: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
por: Mi, Liang, et al.
Publicado: (2026)
por: Mi, Liang, et al.
Publicado: (2026)
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
por: Liu, Jianmin, et al.
Publicado: (2025)
por: Liu, Jianmin, et al.
Publicado: (2025)
FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
por: Li, Rukuo, et al.
Publicado: (2025)
por: Li, Rukuo, et al.
Publicado: (2025)
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
por: Ding, Chuntao, et al.
Publicado: (2024)
por: Ding, Chuntao, et al.
Publicado: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
por: Wang, Wenfeng, et al.
Publicado: (2026)
por: Wang, Wenfeng, et al.
Publicado: (2026)
Federated LoRA with Sparse Communication
por: Kuo, Kevin, et al.
Publicado: (2024)
por: Kuo, Kevin, et al.
Publicado: (2024)
pFedLoRA: Model-Heterogeneous Personalized Federated Learning with LoRA Tuning
por: Yi, Liping, et al.
Publicado: (2023)
por: Yi, Liping, et al.
Publicado: (2023)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
por: Toniolo, Alessio Ricci, et al.
Publicado: (2026)
por: Toniolo, Alessio Ricci, et al.
Publicado: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
por: Zhao, Zhan, et al.
Publicado: (2026)
por: Zhao, Zhan, et al.
Publicado: (2026)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
por: Lee, Sanghyeon, et al.
Publicado: (2025)
por: Lee, Sanghyeon, et al.
Publicado: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
por: Hu, Cunchen, et al.
Publicado: (2024)
por: Hu, Cunchen, et al.
Publicado: (2024)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
por: Chen, Shuangyi, et al.
Publicado: (2024)
por: Chen, Shuangyi, et al.
Publicado: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
por: Shen, Zheyu, et al.
Publicado: (2025)
por: Shen, Zheyu, et al.
Publicado: (2025)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
por: Ni, Yinan, et al.
Publicado: (2025)
por: Ni, Yinan, et al.
Publicado: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
por: Zhu, Zhanda, et al.
Publicado: (2025)
por: Zhu, Zhanda, et al.
Publicado: (2025)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
por: Sui, Yifan, et al.
Publicado: (2025)
por: Sui, Yifan, et al.
Publicado: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
por: Meng, William, et al.
Publicado: (2025)
por: Meng, William, et al.
Publicado: (2025)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
por: Zuo, Jingwei, et al.
Publicado: (2026)
por: Zuo, Jingwei, et al.
Publicado: (2026)
Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?
por: Wang, Yatong, et al.
Publicado: (2026)
por: Wang, Yatong, et al.
Publicado: (2026)
Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation Models
por: Cho, Yae Jee, et al.
Publicado: (2024)
por: Cho, Yae Jee, et al.
Publicado: (2024)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
por: Stepanek, Lukas
Publicado: (2026)
por: Stepanek, Lukas
Publicado: (2026)
Ejemplares similares
-
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024) -
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
por: Wang, Shao, et al.
Publicado: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
por: Sheng, Ying, et al.
Publicado: (2023) -
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
por: Chen, Hongyu, et al.
Publicado: (2026) -
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)