C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Shengye, Xu, Minxian, Zhang, Zuowei, Gao, Chengxi, Zeng, Fansong, Ding, Yu, Ye, Kejiang, Xu, Chengzhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
C‐Koordinator: Interference‐Aware Management for Large‐Scale and Co‐Located Microservice Clusters
by: Shengye Song, et al.
Published: (2026)
by: Shengye Song, et al.
Published: (2026)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
by: Hu, Kan, et al.
Published: (2024)
by: Hu, Kan, et al.
Published: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)
by: Tang, Lujie, et al.
Published: (2024)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
by: Bai, Haoyu, et al.
Published: (2024)
by: Bai, Haoyu, et al.
Published: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
by: Zheng, Wanyi, et al.
Published: (2025)
by: Zheng, Wanyi, et al.
Published: (2025)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
by: Song, Shengye, et al.
Published: (2025)
by: Song, Shengye, et al.
Published: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
by: Hu, Kan, et al.
Published: (2024)
by: Hu, Kan, et al.
Published: (2024)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
by: Wu, Jingfeng, et al.
Published: (2024)
by: Wu, Jingfeng, et al.
Published: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Cloud Native System for LLM Inference Serving
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
by: He, Yiyuan, et al.
Published: (2024)
by: He, Yiyuan, et al.
Published: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
by: Liao, Junhan, et al.
Published: (2025)
by: Liao, Junhan, et al.
Published: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
by: Lin, Yanying, et al.
Published: (2025)
by: Lin, Yanying, et al.
Published: (2025)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
by: Jia, Haojie, et al.
Published: (2025)
by: Jia, Haojie, et al.
Published: (2025)
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
by: Yang, Dingyu, et al.
Published: (2024)
by: Yang, Dingyu, et al.
Published: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
by: He, Yiyuan, et al.
Published: (2025)
by: He, Yiyuan, et al.
Published: (2025)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
by: Bai, Haoyu, et al.
Published: (2026)
by: Bai, Haoyu, et al.
Published: (2026)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
by: Xu, Minxian, et al.
Published: (2026)
by: Xu, Minxian, et al.
Published: (2026)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)
by: Hua, Qin, et al.
Published: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
Energy-aware Distributed Microservice Request Placement at the Edge
by: Toczé, Klervie, et al.
Published: (2024)
by: Toczé, Klervie, et al.
Published: (2024)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
by: Yang, Dingyu, et al.
Published: (2026)
by: Yang, Dingyu, et al.
Published: (2026)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
by: Wu, Siyu, et al.
Published: (2025)
by: Wu, Siyu, et al.
Published: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
by: Mo, Zizhao, et al.
Published: (2026)
by: Mo, Zizhao, et al.
Published: (2026)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
by: Ahmad, Hussain, et al.
Published: (2025)
by: Ahmad, Hussain, et al.
Published: (2025)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
by: Duan, Jiaang, et al.
Published: (2025)
by: Duan, Jiaang, et al.
Published: (2025)
SpotKube: Cost-Optimal Microservices Deployment with Cluster Autoscaling and Spot Pricing
by: Edirisinghe, Dasith, et al.
Published: (2024)
by: Edirisinghe, Dasith, et al.
Published: (2024)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
by: Sun, Yifan, et al.
Published: (2026)
by: Sun, Yifan, et al.
Published: (2026)
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
by: Bai, Xu, et al.
Published: (2025)
by: Bai, Xu, et al.
Published: (2025)
Adaptive Management of Microservices in Dynamic Computing Environments: A Taxonomy and Future Directions
by: Chen, Ming, et al.
Published: (2026)
by: Chen, Ming, et al.
Published: (2026)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
by: Zhang, Ping, et al.
Published: (2024)
by: Zhang, Ping, et al.
Published: (2024)
Edge AI: A Taxonomy, Systematic Review and Future Directions
by: Gill, Sukhpal Singh, et al.
Published: (2024)
by: Gill, Sukhpal Singh, et al.
Published: (2024)
Collaborative Evolution of Intelligent Agents in Large-Scale Microservice Systems
by: Li, Yilin, et al.
Published: (2025)
by: Li, Yilin, et al.
Published: (2025)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
Metric Criticality Identification for Cloud Microservices
by: Singal, Akanksha, et al.
Published: (2025)
by: Singal, Akanksha, et al.
Published: (2025)
Similar Items
-
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
by: Li, Xiang, et al.
Published: (2024) -
C‐Koordinator: Interference‐Aware Management for Large‐Scale and Co‐Located Microservice Clusters
by: Shengye Song, et al.
Published: (2026) -
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
by: Hu, Kan, et al.
Published: (2024) -
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
by: Xu, Minxian, et al.
Published: (2025) -
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)