SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Tian, Mao, Ziming, Kerney, Jamison, Jackson, Ethan J., Li, Zhifei, Xing, Jiarong, Shenker, Scott, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
by: Li, Zhifei, et al.
Published: (2026)
by: Li, Zhifei, et al.
Published: (2026)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
by: Griggs, Tyler, et al.
Published: (2026)
by: Griggs, Tyler, et al.
Published: (2026)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
by: Wang, Yingping, et al.
Published: (2026)
by: Wang, Yingping, et al.
Published: (2026)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
by: Bournias, Ilias, et al.
Published: (2024)
by: Bournias, Ilias, et al.
Published: (2024)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
by: Zhang, Wenxin, et al.
Published: (2024)
by: Zhang, Wenxin, et al.
Published: (2024)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
by: Liu, Xueshen, et al.
Published: (2026)
by: Liu, Xueshen, et al.
Published: (2026)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
Morpheus: Lightweight RTT Prediction for Performance-Aware Load Balancing
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
Review of Hybrid Load Balancing Algorithms in Cloud Computing Environment
by: Ijeoma, Chukwuneke Chiamaka, et al.
Published: (2022)
by: Ijeoma, Chukwuneke Chiamaka, et al.
Published: (2022)
Load Balanced Parallel Node Generation for Meshless Numerical Methods
by: Vehovar, Jon, et al.
Published: (2026)
by: Vehovar, Jon, et al.
Published: (2026)
Fine-grained MoE Load Balancing with Linear Programming
by: Zhao, Chenqi, et al.
Published: (2025)
by: Zhao, Chenqi, et al.
Published: (2025)
A Communication- and Memory-Aware Model for Load Balancing Tasks
by: Lifflander, Jonathan, et al.
Published: (2024)
by: Lifflander, Jonathan, et al.
Published: (2024)
UCCL-EP: Portable Expert-Parallel Communication
by: Mao, Ziming, et al.
Published: (2025)
by: Mao, Ziming, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
Locality-aware Fair Scheduling in LLM Serving
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
by: Ma, Shuang, et al.
Published: (2026)
by: Ma, Shuang, et al.
Published: (2026)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Load Balancing in Strongly Inhomogeneous Simulations -- a Vlasiator Case Study
by: Kotipalo, Leo, et al.
Published: (2025)
by: Kotipalo, Leo, et al.
Published: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
Inference Load-Aware Orchestration for Hierarchical Federated Learning
by: Lackinger, Anna, et al.
Published: (2024)
by: Lackinger, Anna, et al.
Published: (2024)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
by: Zhao, Yilong, et al.
Published: (2026)
by: Zhao, Yilong, et al.
Published: (2026)
TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
by: Zhao, Yiwei, et al.
Published: (2025)
by: Zhao, Yiwei, et al.
Published: (2025)
LB4OMP: A Dynamic Load Balancing Library for Multithreaded Applications
by: Korndörfer, Jonas H. Müller, et al.
Published: (2021)
by: Korndörfer, Jonas H. Müller, et al.
Published: (2021)
QEdgeProxy: QoS-Aware Load Balancing for IoT Services in the Computing Continuum
by: Čilić, Ivan, et al.
Published: (2024)
by: Čilić, Ivan, et al.
Published: (2024)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
by: Sandholm, Thomas, et al.
Published: (2025)
by: Sandholm, Thomas, et al.
Published: (2025)
(Almost) Perfect Discrete Iterative Load Balancing
by: Berenbrink, Petra, et al.
Published: (2025)
by: Berenbrink, Petra, et al.
Published: (2025)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
by: Yuan, Ying, et al.
Published: (2026)
by: Yuan, Ying, et al.
Published: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
An Analytical Overview Of Virtual Machine Load Balancing Scheduling Algorithms with their Comparative Case Study
by: Vaidya, Priyank, et al.
Published: (2025)
by: Vaidya, Priyank, et al.
Published: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
Exploring Dynamic Load Balancing Algorithms for Block-Structured Mesh-and-Particle Simulations in AMReX
by: Nanda, Amitash, et al.
Published: (2025)
by: Nanda, Amitash, et al.
Published: (2025)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
by: Qi, Shuyao, et al.
Published: (2026)
by: Qi, Shuyao, et al.
Published: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Similar Items
-
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
by: Li, Zhifei, et al.
Published: (2026) -
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024) -
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024) -
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024) -
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
by: Liu, Shu, et al.
Published: (2025)