A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
Fuente:
arXiv
Guardado en:
| Autores principales: | Jeon, Beomyeol, Wang, Chen, Arroyo, Diana, Youssef, Alaa, Gupta, Indranil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SLO-Aware Scheduling for Large Language Model Inferences
por: Huang, Jinqi, et al.
Publicado: (2025)
por: Huang, Jinqi, et al.
Publicado: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
por: Wang, Qipeng
Publicado: (2026)
por: Wang, Qipeng
Publicado: (2026)
Syndeo: Portable Ray Clusters with Secure Containerization
por: Li, William, et al.
Publicado: (2024)
por: Li, William, et al.
Publicado: (2024)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
por: Zhang, Zongshun, et al.
Publicado: (2025)
por: Zhang, Zongshun, et al.
Publicado: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
por: K., Prashanthi S., et al.
Publicado: (2023)
por: K., Prashanthi S., et al.
Publicado: (2023)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
por: Chen, Jiabin, et al.
Publicado: (2024)
por: Chen, Jiabin, et al.
Publicado: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025)
por: Hu, Jianmin, et al.
Publicado: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
por: Cheng, Ke, et al.
Publicado: (2024)
por: Cheng, Ke, et al.
Publicado: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
por: Nie, Chengyi, et al.
Publicado: (2024)
por: Nie, Chengyi, et al.
Publicado: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
por: Sedlak, Boris, et al.
Publicado: (2024)
por: Sedlak, Boris, et al.
Publicado: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
por: Ma, Chenxiang, et al.
Publicado: (2025)
por: Ma, Chenxiang, et al.
Publicado: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
por: Chow, Will
Publicado: (2025)
por: Chow, Will
Publicado: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
por: Gao, Wei, et al.
Publicado: (2026)
por: Gao, Wei, et al.
Publicado: (2026)
Adaptive Resource Allocation for Workflow Containerization on Kubernetes
por: Shan, Chenggang, et al.
Publicado: (2023)
por: Shan, Chenggang, et al.
Publicado: (2023)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
por: Jain, Rutwik, et al.
Publicado: (2024)
por: Jain, Rutwik, et al.
Publicado: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
por: Hu, Tiancheng, et al.
Publicado: (2026)
por: Hu, Tiancheng, et al.
Publicado: (2026)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
por: Medeiros, Daniel, et al.
Publicado: (2024)
por: Medeiros, Daniel, et al.
Publicado: (2024)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
por: Xu, Ao, et al.
Publicado: (2025)
por: Xu, Ao, et al.
Publicado: (2025)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
por: Medeiros, Daniel, et al.
Publicado: (2024)
por: Medeiros, Daniel, et al.
Publicado: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
por: Gu, Jianfeng, et al.
Publicado: (2025)
por: Gu, Jianfeng, et al.
Publicado: (2025)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
por: Punniyamoorthy, Vinoth, et al.
Publicado: (2025)
por: Punniyamoorthy, Vinoth, et al.
Publicado: (2025)
ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
por: Medeiros, Daniel, et al.
Publicado: (2025)
por: Medeiros, Daniel, et al.
Publicado: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
por: Rajashekar, Kolichala, et al.
Publicado: (2025)
por: Rajashekar, Kolichala, et al.
Publicado: (2025)
Containerization in Multi-Cloud Environment: Roles, Strategies, Challenges, and Solutions for Effective Implementation
por: Waseem, Muhammad, et al.
Publicado: (2024)
por: Waseem, Muhammad, et al.
Publicado: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
por: Qi, S., et al.
Publicado: (2024)
por: Qi, S., et al.
Publicado: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
por: Banegas-Luna, Antonio Jesús, et al.
Publicado: (2025)
por: Banegas-Luna, Antonio Jesús, et al.
Publicado: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
por: Xuan, Mo, et al.
Publicado: (2025)
por: Xuan, Mo, et al.
Publicado: (2025)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
por: Shan, Chenggang, et al.
Publicado: (2024)
por: Shan, Chenggang, et al.
Publicado: (2024)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
por: Li, Luchang, et al.
Publicado: (2026)
por: Li, Luchang, et al.
Publicado: (2026)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
por: Peng, Haosong, et al.
Publicado: (2024)
por: Peng, Haosong, et al.
Publicado: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
por: Shen, Haiying, et al.
Publicado: (2024)
por: Shen, Haiying, et al.
Publicado: (2024)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
por: Sun, Desen, et al.
Publicado: (2025)
por: Sun, Desen, et al.
Publicado: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
por: Huang, Shaoyuan, et al.
Publicado: (2026)
por: Huang, Shaoyuan, et al.
Publicado: (2026)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
por: Yu, Jiahuan, et al.
Publicado: (2026)
por: Yu, Jiahuan, et al.
Publicado: (2026)
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
por: kumar, Ramakant
Publicado: (2026)
por: kumar, Ramakant
Publicado: (2026)
EcoServe: Designing Carbon-Aware AI Inference Systems
por: Li, Yueying, et al.
Publicado: (2025)
por: Li, Yueying, et al.
Publicado: (2025)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
por: Hu, Kan, et al.
Publicado: (2024)
por: Hu, Kan, et al.
Publicado: (2024)
ProvDeploy: Provenance-oriented Containerization of High Performance Computing Scientific Workflows
por: Kunstmann, Liliane, et al.
Publicado: (2024)
por: Kunstmann, Liliane, et al.
Publicado: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
por: Hu, Kan, et al.
Publicado: (2024)
por: Hu, Kan, et al.
Publicado: (2024)
Ejemplares similares
-
SLO-Aware Scheduling for Large Language Model Inferences
por: Huang, Jinqi, et al.
Publicado: (2025) -
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
por: Wang, Qipeng
Publicado: (2026) -
Syndeo: Portable Ray Clusters with Secure Containerization
por: Li, William, et al.
Publicado: (2024) -
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
por: Zhang, Zongshun, et al.
Publicado: (2025) -
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
por: K., Prashanthi S., et al.
Publicado: (2023)