DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Ying, Zuo, Pengfei, Wang, Bo, Chen, Zhangyu, Tan, Zhipeng, Yu, Zhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
di: Liang, Yunkai, et al.
Pubblicazione: (2025)
di: Liang, Yunkai, et al.
Pubblicazione: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
di: Bu, Tianci, et al.
Pubblicazione: (2026)
di: Bu, Tianci, et al.
Pubblicazione: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
di: Nian, Sean, et al.
Pubblicazione: (2026)
di: Nian, Sean, et al.
Pubblicazione: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
di: Da, Wei, et al.
Pubblicazione: (2025)
di: Da, Wei, et al.
Pubblicazione: (2025)
InstCache: A Predictive Cache for LLM Serving
di: Zou, Longwei, et al.
Pubblicazione: (2024)
di: Zou, Longwei, et al.
Pubblicazione: (2024)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
di: Ai, Xin, et al.
Pubblicazione: (2024)
di: Ai, Xin, et al.
Pubblicazione: (2024)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
di: Zhang, Wenxin, et al.
Pubblicazione: (2024)
di: Zhang, Wenxin, et al.
Pubblicazione: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
di: Gandhi, Rohan, et al.
Pubblicazione: (2024)
di: Gandhi, Rohan, et al.
Pubblicazione: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
di: Wang, Wenfeng, et al.
Pubblicazione: (2026)
di: Wang, Wenfeng, et al.
Pubblicazione: (2026)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
di: Bournias, Ilias, et al.
Pubblicazione: (2024)
di: Bournias, Ilias, et al.
Pubblicazione: (2024)
A Universal Load Balancing Principle and Its Application to Large Language Model Serving
di: Chen, Zixi, et al.
Pubblicazione: (2026)
di: Chen, Zixi, et al.
Pubblicazione: (2026)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025)
di: Xie, Zhiqiang, et al.
Pubblicazione: (2025)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
di: Xu, Jiale, et al.
Pubblicazione: (2025)
di: Xu, Jiale, et al.
Pubblicazione: (2025)
TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
di: Zhao, Yiwei, et al.
Pubblicazione: (2025)
di: Zhao, Yiwei, et al.
Pubblicazione: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
di: Yoshimura, Takeshi, et al.
Pubblicazione: (2026)
di: Yoshimura, Takeshi, et al.
Pubblicazione: (2026)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
Fine-grained MoE Load Balancing with Linear Programming
di: Zhao, Chenqi, et al.
Pubblicazione: (2025)
di: Zhao, Chenqi, et al.
Pubblicazione: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
di: Xia, Tian, et al.
Pubblicazione: (2025)
di: Xia, Tian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025) -
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
di: Liang, Yunkai, et al.
Pubblicazione: (2025) -
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025) -
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
di: Bu, Tianci, et al.
Pubblicazione: (2026) -
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
di: Yuan, Yitao, et al.
Pubblicazione: (2025)