Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Mei, Yixuan, Li, Zikun, Chen, Zixuan, Pan, Shiqi, Wu, Mengdi, Miao, Xupeng, Jia, Zhihao, Rashmi, K. V. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
por: Mei, Yixuan, et al.
Publicado: (2024)
por: Mei, Yixuan, et al.
Publicado: (2024)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
por: Jiang, Youhe, et al.
Publicado: (2026)
por: Jiang, Youhe, et al.
Publicado: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
por: He, Xuan, et al.
Publicado: (2025)
por: He, Xuan, et al.
Publicado: (2025)
Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUs (Extended Version)
por: Xu, Mingkuan, et al.
Publicado: (2024)
por: Xu, Mingkuan, et al.
Publicado: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
por: Oliaro, Gabriele, et al.
Publicado: (2024)
por: Oliaro, Gabriele, et al.
Publicado: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026)
por: Du, Boxiao, et al.
Publicado: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
por: Li, Zikun, et al.
Publicado: (2025)
por: Li, Zikun, et al.
Publicado: (2025)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
por: Xia, Yifei, et al.
Publicado: (2025)
por: Xia, Yifei, et al.
Publicado: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
por: He, Wenhao, et al.
Publicado: (2026)
por: He, Wenhao, et al.
Publicado: (2026)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
por: Kim, Kihyun, et al.
Publicado: (2025)
por: Kim, Kihyun, et al.
Publicado: (2025)
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
por: Lu, Runyu, et al.
Publicado: (2025)
por: Lu, Runyu, et al.
Publicado: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
por: Shi, Tianyao, et al.
Publicado: (2024)
por: Shi, Tianyao, et al.
Publicado: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
por: Peng, You, et al.
Publicado: (2026)
por: Peng, You, et al.
Publicado: (2026)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
por: Miao, Xupeng, et al.
Publicado: (2023)
por: Miao, Xupeng, et al.
Publicado: (2023)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
por: Lou, Chiheng, et al.
Publicado: (2025)
por: Lou, Chiheng, et al.
Publicado: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
por: Zhang, Chen, et al.
Publicado: (2025)
por: Zhang, Chen, et al.
Publicado: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
por: Du, Jiangsu, et al.
Publicado: (2025)
por: Du, Jiangsu, et al.
Publicado: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
por: Gao, Wei, et al.
Publicado: (2026)
por: Gao, Wei, et al.
Publicado: (2026)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
por: Ye, Fanjiang, et al.
Publicado: (2026)
por: Ye, Fanjiang, et al.
Publicado: (2026)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
por: Duan, Jiangfei, et al.
Publicado: (2024)
por: Duan, Jiangfei, et al.
Publicado: (2024)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
por: Yan, Ran, et al.
Publicado: (2025)
por: Yan, Ran, et al.
Publicado: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2025)
por: Xu, Jiale, et al.
Publicado: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
por: Wilkins, Grant, et al.
Publicado: (2024)
por: Wilkins, Grant, et al.
Publicado: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
por: Chen, Xing, et al.
Publicado: (2025)
por: Chen, Xing, et al.
Publicado: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
por: Zhao, Zhixin, et al.
Publicado: (2024)
por: Zhao, Zhixin, et al.
Publicado: (2024)
A System for Microserving of LLMs
por: Jin, Hongyi, et al.
Publicado: (2024)
por: Jin, Hongyi, et al.
Publicado: (2024)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
por: Griggs, Tyler, et al.
Publicado: (2024)
por: Griggs, Tyler, et al.
Publicado: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
por: Wang, Weiye, et al.
Publicado: (2026)
por: Wang, Weiye, et al.
Publicado: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
por: He, Yiyuan, et al.
Publicado: (2024)
por: He, Yiyuan, et al.
Publicado: (2024)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
por: Kim, Heehoon, et al.
Publicado: (2026)
por: Kim, Heehoon, et al.
Publicado: (2026)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
por: Zhan, Huiyou, et al.
Publicado: (2025)
por: Zhan, Huiyou, et al.
Publicado: (2025)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
por: Liu, Lian, et al.
Publicado: (2026)
por: Liu, Lian, et al.
Publicado: (2026)
Ejemplares similares
-
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
por: Mei, Yixuan, et al.
Publicado: (2024) -
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025) -
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
por: Jiang, Youhe, et al.
Publicado: (2026) -
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
por: He, Xuan, et al.
Publicado: (2025) -
Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUs (Extended Version)
por: Xu, Mingkuan, et al.
Publicado: (2024)