Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted Microservices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zibo, Li, Pinghe, Liang, Chieh-Jan Mike, Wu, Feng, Yan, Francis Y. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025)
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
Decouple and Decompose: Scaling Resource Allocation with DeDe
von: Xu, Zhiying, et al.
Veröffentlicht: (2024)
von: Xu, Zhiying, et al.
Veröffentlicht: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
von: Sedlak, Boris, et al.
Veröffentlicht: (2024)
von: Sedlak, Boris, et al.
Veröffentlicht: (2024)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
von: Sun, Desen, et al.
Veröffentlicht: (2025)
von: Sun, Desen, et al.
Veröffentlicht: (2025)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
von: Xuan, Mo, et al.
Veröffentlicht: (2025)
von: Xuan, Mo, et al.
Veröffentlicht: (2025)
Analytically-Driven Resource Management for Cloud-Native Microservices
von: Zhang, Yanqi, et al.
Veröffentlicht: (2024)
von: Zhang, Yanqi, et al.
Veröffentlicht: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
von: Legler, Julian, et al.
Veröffentlicht: (2025)
von: Legler, Julian, et al.
Veröffentlicht: (2025)
Autonomous Resource Management in Microservice Systems via Reinforcement Learning
von: Zou, Yujun, et al.
Veröffentlicht: (2025)
von: Zou, Yujun, et al.
Veröffentlicht: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Artifact for Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
von: Legler, Julian
Veröffentlicht: (2026)
von: Legler, Julian
Veröffentlicht: (2026)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
von: Yang, Dingyu, et al.
Veröffentlicht: (2026)
von: Yang, Dingyu, et al.
Veröffentlicht: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
von: Chow, Will
Veröffentlicht: (2025)
von: Chow, Will
Veröffentlicht: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Adaptive Management of Microservices in Dynamic Computing Environments: A Taxonomy and Future Directions
von: Chen, Ming, et al.
Veröffentlicht: (2026)
von: Chen, Ming, et al.
Veröffentlicht: (2026)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
von: Song, Shengye, et al.
Veröffentlicht: (2025)
von: Song, Shengye, et al.
Veröffentlicht: (2025)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
von: Lu, Baotong, et al.
Veröffentlicht: (2024)
von: Lu, Baotong, et al.
Veröffentlicht: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge Systems
von: Wen, Yangyang, et al.
Veröffentlicht: (2025)
von: Wen, Yangyang, et al.
Veröffentlicht: (2025)
Metric Criticality Identification for Cloud Microservices
von: Singal, Akanksha, et al.
Veröffentlicht: (2025)
von: Singal, Akanksha, et al.
Veröffentlicht: (2025)
From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
von: Cui, Xingqi, et al.
Veröffentlicht: (2025)
von: Cui, Xingqi, et al.
Veröffentlicht: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
Large-scale Neural Network Quantum States for ab initio Quantum Chemistry Simulations on Fugaku
von: Xu, Hongtao, et al.
Veröffentlicht: (2025)
von: Xu, Hongtao, et al.
Veröffentlicht: (2025)
Signalling Health for Improved Kubernetes Microservice Availability
von: Roberts, Jacob, et al.
Veröffentlicht: (2025)
von: Roberts, Jacob, et al.
Veröffentlicht: (2025)
Self-adaptive, Requirements-driven Autoscaling of Microservices
von: Nunes, João Paulo Karol Santos, et al.
Veröffentlicht: (2024)
von: Nunes, João Paulo Karol Santos, et al.
Veröffentlicht: (2024)
NotNets: Accelerating Microservices by Bypassing the Network
von: Alvaro, Peter, et al.
Veröffentlicht: (2024)
von: Alvaro, Peter, et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
von: Jeon, Beomyeol, et al.
Veröffentlicht: (2024)
von: Jeon, Beomyeol, et al.
Veröffentlicht: (2024)
Energy-aware Distributed Microservice Request Placement at the Edge
von: Toczé, Klervie, et al.
Veröffentlicht: (2024)
von: Toczé, Klervie, et al.
Veröffentlicht: (2024)
Energy Metrics for Edge Microservice Request Placement Strategies
von: Toczé, Klervie, et al.
Veröffentlicht: (2025)
von: Toczé, Klervie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024) -
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024) -
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025) -
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024) -
Decouple and Decompose: Scaling Resource Allocation with DeDe
von: Xu, Zhiying, et al.
Veröffentlicht: (2024)