An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xiang, Wen, Linfeng, Xu, Minxian, Ye, Kejiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
di: Song, Shengye, et al.
Pubblicazione: (2025)
di: Song, Shengye, et al.
Pubblicazione: (2025)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
di: Song, Shengye, et al.
Pubblicazione: (2025)
di: Song, Shengye, et al.
Pubblicazione: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
di: Bai, Haoyu, et al.
Pubblicazione: (2024)
di: Bai, Haoyu, et al.
Pubblicazione: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
di: Jia, Haojie, et al.
Pubblicazione: (2025)
di: Jia, Haojie, et al.
Pubblicazione: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
di: Tang, Lujie, et al.
Pubblicazione: (2024)
di: Tang, Lujie, et al.
Pubblicazione: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Evaluating Container Orchestration for Neuromorphic Workloads in Virtual Edge Environments
di: Pham, Huyen, et al.
Pubblicazione: (2026)
di: Pham, Huyen, et al.
Pubblicazione: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
di: Lin, Yanying, et al.
Pubblicazione: (2025)
di: Lin, Yanying, et al.
Pubblicazione: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
di: Xuan, Mo, et al.
Pubblicazione: (2025)
di: Xuan, Mo, et al.
Pubblicazione: (2025)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
di: Bai, Haoyu, et al.
Pubblicazione: (2026)
di: Bai, Haoyu, et al.
Pubblicazione: (2026)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
di: Zhang, Ping, et al.
Pubblicazione: (2024)
di: Zhang, Ping, et al.
Pubblicazione: (2024)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
di: Sun, Yifan, et al.
Pubblicazione: (2026)
di: Sun, Yifan, et al.
Pubblicazione: (2026)
Resource Slicing through Intelligent Orchestration of Energy-aware IoT services in Edge-Cloud Continuum
di: Shahid, Hafiz Faheem, et al.
Pubblicazione: (2024)
di: Shahid, Hafiz Faheem, et al.
Pubblicazione: (2024)
Edge AI: A Taxonomy, Systematic Review and Future Directions
di: Gill, Sukhpal Singh, et al.
Pubblicazione: (2024)
di: Gill, Sukhpal Singh, et al.
Pubblicazione: (2024)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
di: Xu, Minxian, et al.
Pubblicazione: (2026)
di: Xu, Minxian, et al.
Pubblicazione: (2026)
LRScheduler: A Layer-aware and Resource-adaptive Container Scheduler in Edge Computing
di: Tang, Zhiqing, et al.
Pubblicazione: (2025)
di: Tang, Zhiqing, et al.
Pubblicazione: (2025)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
di: Biswas, Anish, et al.
Pubblicazione: (2026)
di: Biswas, Anish, et al.
Pubblicazione: (2026)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
di: Li, Maoliang, et al.
Pubblicazione: (2026)
di: Li, Maoliang, et al.
Pubblicazione: (2026)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration
di: Baghel, Himanshu Singh
Pubblicazione: (2026)
di: Baghel, Himanshu Singh
Pubblicazione: (2026)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
di: Wu, Siyu, et al.
Pubblicazione: (2025)
di: Wu, Siyu, et al.
Pubblicazione: (2025)
CIR: Lightweight Container Image for Cross-Platform Deployment
di: Li, Fengzhi, et al.
Pubblicazione: (2026)
di: Li, Fengzhi, et al.
Pubblicazione: (2026)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
di: Saba, Issa, et al.
Pubblicazione: (2024)
di: Saba, Issa, et al.
Pubblicazione: (2024)
iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration
di: Guan, Wen, et al.
Pubblicazione: (2025)
di: Guan, Wen, et al.
Pubblicazione: (2025)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
di: Copik, Marcin, et al.
Pubblicazione: (2025)
di: Copik, Marcin, et al.
Pubblicazione: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
di: Liang, Antian, et al.
Pubblicazione: (2025)
di: Liang, Antian, et al.
Pubblicazione: (2025)
Orchestrating the Execution of Serverless Functions in Hybrid Clouds
di: Peri, Aristotelis, et al.
Pubblicazione: (2024)
di: Peri, Aristotelis, et al.
Pubblicazione: (2024)
Documenti analoghi
-
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
di: Song, Shengye, et al.
Pubblicazione: (2025) -
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
di: Wen, Linfeng, et al.
Pubblicazione: (2024) -
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
di: Hu, Kan, et al.
Pubblicazione: (2024) -
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
di: Song, Shengye, et al.
Pubblicazione: (2025) -
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
di: Bai, Haoyu, et al.
Pubblicazione: (2024)