Predictable LLM Serving on GPU Clusters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Darzi, Erfan, Bharadwaj, Shreeanant, Balija, Sree Bhargavi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
von: Kong, Z. Jonny, et al.
Veröffentlicht: (2025)
von: Kong, Z. Jonny, et al.
Veröffentlicht: (2025)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Fast State Restoration in LLM Serving with HCache
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
von: Mo, Zizhao, et al.
Veröffentlicht: (2024)
von: Mo, Zizhao, et al.
Veröffentlicht: (2024)
DeepOps & SLURM: Your GPU Cluster Guide
von: Majee, Arindam
Veröffentlicht: (2024)
von: Majee, Arindam
Veröffentlicht: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
InstCache: A Predictive Cache for LLM Serving
von: Zou, Longwei, et al.
Veröffentlicht: (2024)
von: Zou, Longwei, et al.
Veröffentlicht: (2024)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
von: Darzi, Erfan, et al.
Veröffentlicht: (2025) -
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
von: Mo, Zizhao, et al.
Veröffentlicht: (2025) -
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
von: Lin, Zejia, et al.
Veröffentlicht: (2025) -
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026) -
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
von: Kong, Z. Jonny, et al.
Veröffentlicht: (2025)