ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Shao, Ren, Rui, Gui, Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
di: Li, Allison, et al.
Pubblicazione: (2025)
di: Li, Allison, et al.
Pubblicazione: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
di: Sheng, Ying, et al.
Pubblicazione: (2023)
di: Sheng, Ying, et al.
Pubblicazione: (2023)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
di: Guo, Yipin, et al.
Pubblicazione: (2026)
di: Guo, Yipin, et al.
Pubblicazione: (2026)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
di: Patel, Ishan, et al.
Pubblicazione: (2026)
di: Patel, Ishan, et al.
Pubblicazione: (2026)
Federated LoRA with Sparse Communication
di: Kuo, Kevin, et al.
Pubblicazione: (2024)
di: Kuo, Kevin, et al.
Pubblicazione: (2024)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
di: Xiong, Yi, et al.
Pubblicazione: (2024)
di: Xiong, Yi, et al.
Pubblicazione: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
di: Mi, Liang, et al.
Pubblicazione: (2026)
di: Mi, Liang, et al.
Pubblicazione: (2026)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
pFedLoRA: Model-Heterogeneous Personalized Federated Learning with LoRA Tuning
di: Yi, Liping, et al.
Pubblicazione: (2023)
di: Yi, Liping, et al.
Pubblicazione: (2023)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
di: Sui, Yifan, et al.
Pubblicazione: (2025)
di: Sui, Yifan, et al.
Pubblicazione: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
di: Brüel-Gabrielsson, Rickard, et al.
Pubblicazione: (2024)
di: Brüel-Gabrielsson, Rickard, et al.
Pubblicazione: (2024)
Leyline: KV Cache Directives for Agentic Inference
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
di: Chen, Shuangyi, et al.
Pubblicazione: (2024)
di: Chen, Shuangyi, et al.
Pubblicazione: (2024)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
di: Zuo, Jingwei, et al.
Pubblicazione: (2026)
di: Zuo, Jingwei, et al.
Pubblicazione: (2026)
Fed-pilot: Optimizing LoRA Allocation for Efficient Federated Fine-Tuning with Heterogeneous Clients
di: Zhang, Zikai, et al.
Pubblicazione: (2024)
di: Zhang, Zikai, et al.
Pubblicazione: (2024)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
di: Nian, Sean, et al.
Pubblicazione: (2026)
di: Nian, Sean, et al.
Pubblicazione: (2026)
FedRPCA: Enhancing Federated LoRA Aggregation Using Robust PCA
di: Jhunjhunwala, Divyansh, et al.
Pubblicazione: (2025)
di: Jhunjhunwala, Divyansh, et al.
Pubblicazione: (2025)
Stabilizing Decentralized Federated Fine-Tuning via Topology-Aware Alternating LoRA
di: Wang, Xiaoyu, et al.
Pubblicazione: (2026)
di: Wang, Xiaoyu, et al.
Pubblicazione: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
di: Qianli, Liu, et al.
Pubblicazione: (2025)
di: Qianli, Liu, et al.
Pubblicazione: (2025)
Fed-HeLLo: Efficient Federated Foundation Model Fine-Tuning with Heterogeneous LoRA Allocation
di: Zhang, Zikai, et al.
Pubblicazione: (2025)
di: Zhang, Zikai, et al.
Pubblicazione: (2025)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
di: Lee, Wonbeom, et al.
Pubblicazione: (2024)
di: Lee, Wonbeom, et al.
Pubblicazione: (2024)
Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation Models
di: Cho, Yae Jee, et al.
Pubblicazione: (2024)
di: Cho, Yae Jee, et al.
Pubblicazione: (2024)
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
di: Zhang, Yanqi, et al.
Pubblicazione: (2024)
di: Zhang, Yanqi, et al.
Pubblicazione: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
di: Jin, Yibo, et al.
Pubblicazione: (2024)
di: Jin, Yibo, et al.
Pubblicazione: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
AutoRank: MCDA Based Rank Personalization for LoRA-Enabled Distributed Learning
di: Chen, Shuaijun, et al.
Pubblicazione: (2024)
di: Chen, Shuaijun, et al.
Pubblicazione: (2024)
RBLA: Rank-Based-LoRA-Aggregation for Fine-tuning Heterogeneous Models in FLaaS
di: Chen, Shuaijun, et al.
Pubblicazione: (2024)
di: Chen, Shuaijun, et al.
Pubblicazione: (2024)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
di: Liu, Zedong, et al.
Pubblicazione: (2026)
di: Liu, Zedong, et al.
Pubblicazione: (2026)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
di: Zhu, Zhanda, et al.
Pubblicazione: (2025)
di: Zhu, Zhanda, et al.
Pubblicazione: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
di: Li, Allison, et al.
Pubblicazione: (2025) -
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026) -
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
di: Sheng, Ying, et al.
Pubblicazione: (2023) -
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)