MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xuan, Mo, yue, Zhang, Weigang, Wu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
par: Peng, Haosong, et autres
Publié: (2024)
par: Peng, Haosong, et autres
Publié: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025)
par: Du, Jiangsu, et autres
Publié: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025)
par: Gu, Jianfeng, et autres
Publié: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
par: Ma, Chenxiang, et autres
Publié: (2025)
par: Ma, Chenxiang, et autres
Publié: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
par: Gao, Wei, et autres
Publié: (2026)
par: Gao, Wei, et autres
Publié: (2026)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
par: Liang, Antian, et autres
Publié: (2025)
par: Liang, Antian, et autres
Publié: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024)
par: Nie, Chengyi, et autres
Publié: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
par: Sedlak, Boris, et autres
Publié: (2024)
par: Sedlak, Boris, et autres
Publié: (2024)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
par: Li, Xiang, et autres
Publié: (2024)
par: Li, Xiang, et autres
Publié: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
par: Zhang, Wei, et autres
Publié: (2025)
par: Zhang, Wei, et autres
Publié: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
par: Shen, Haiying, et autres
Publié: (2024)
par: Shen, Haiying, et autres
Publié: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
par: Duan, Jiaang, et autres
Publié: (2025)
par: Duan, Jiaang, et autres
Publié: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
par: Li, Maoliang, et autres
Publié: (2026)
par: Li, Maoliang, et autres
Publié: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
par: Sun, Desen, et autres
Publié: (2025)
par: Sun, Desen, et autres
Publié: (2025)
Resource Slicing through Intelligent Orchestration of Energy-aware IoT services in Edge-Cloud Continuum
par: Shahid, Hafiz Faheem, et autres
Publié: (2024)
par: Shahid, Hafiz Faheem, et autres
Publié: (2024)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026)
par: Bai, Fengyao, et autres
Publié: (2026)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration
par: Kumar, Satyam, et autres
Publié: (2026)
par: Kumar, Satyam, et autres
Publié: (2026)
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
par: Kulagina, Svetlana, et autres
Publié: (2025)
par: Kulagina, Svetlana, et autres
Publié: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
par: Jeon, Beomyeol, et autres
Publié: (2024)
par: Jeon, Beomyeol, et autres
Publié: (2024)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
par: Mo, Zizhao, et autres
Publié: (2024)
par: Mo, Zizhao, et autres
Publié: (2024)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
par: Saba, Issa, et autres
Publié: (2024)
par: Saba, Issa, et autres
Publié: (2024)
pBeeGees: A Prudent Approach to Certificate-Decoupled BFT Consensus
par: Yang, Kaiji, et autres
Publié: (2025)
par: Yang, Kaiji, et autres
Publié: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
par: Luo, Shuqing, et autres
Publié: (2024)
par: Luo, Shuqing, et autres
Publié: (2024)
Mitigating Artifacts in Pre-quantization Based Scientific Data Compressors with Quantization-aware Interpolation
par: Jiao, Pu, et autres
Publié: (2026)
par: Jiao, Pu, et autres
Publié: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
par: Qi, S., et autres
Publié: (2024)
par: Qi, S., et autres
Publié: (2024)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
par: Zhang, Yue, et autres
Publié: (2025)
par: Zhang, Yue, et autres
Publié: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025)
par: Zhu, Kan, et autres
Publié: (2025)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
par: Punniyamoorthy, Vinoth, et autres
Publié: (2025)
par: Punniyamoorthy, Vinoth, et autres
Publié: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025)
par: Chen, Siyuan, et autres
Publié: (2025)
QoS-aware Scheduling of Periodic Real-time Task Graphs on Heterogeneous Pre-occupied MECs
par: Shankar, Ashutosh, et autres
Publié: (2025)
par: Shankar, Ashutosh, et autres
Publié: (2025)
Documents similaires
-
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
par: Peng, Haosong, et autres
Publié: (2024) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024) -
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026) -
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025) -
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025)