Analytically-Driven Resource Management for Cloud-Native Microservices
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yanqi, Zhou, Zhuangzhuang, Elnikety, Sameh, Delimitrou, Christina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
by: Liang, Mingyu, et al.
Published: (2025)
by: Liang, Mingyu, et al.
Published: (2025)
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
Adaptive AI-based Decentralized Resource Management in the Cloud-Edge Continuum
by: Li, Lanpei, et al.
Published: (2025)
by: Li, Lanpei, et al.
Published: (2025)
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025)
by: Chockalingam, Nachiappan, et al.
Published: (2025)
Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
by: Legler, Julian, et al.
Published: (2025)
by: Legler, Julian, et al.
Published: (2025)
Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review
by: Gu, Yan, et al.
Published: (2025)
by: Gu, Yan, et al.
Published: (2025)
Artifact for Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
by: Legler, Julian
Published: (2026)
by: Legler, Julian
Published: (2026)
Application of Machine Learning Optimization in Cloud Computing Resource Scheduling and Management
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud
by: Wang, Qinlong, et al.
Published: (2023)
by: Wang, Qinlong, et al.
Published: (2023)
Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis
by: Shyam, Gopal Krishna, et al.
Published: (2026)
by: Shyam, Gopal Krishna, et al.
Published: (2026)
Optimized Cloud Resource Allocation Using Genetic Algorithms for Energy Efficiency and QoS Assurance
by: Panggabean, Caroline, et al.
Published: (2025)
by: Panggabean, Caroline, et al.
Published: (2025)
Thousand-GPU Large-Scale Training and Optimization Recipe for AI-Native Cloud Embodied Intelligence Infrastructure
by: Guo, Yongjian, et al.
Published: (2026)
by: Guo, Yongjian, et al.
Published: (2026)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
by: Ahmad, Hussain, et al.
Published: (2025)
by: Ahmad, Hussain, et al.
Published: (2025)
CloudEval-YAML: A Practical Benchmark for Cloud Configuration Generation
by: Xu, Yifei, et al.
Published: (2023)
by: Xu, Yifei, et al.
Published: (2023)
iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
by: Hu, Yi-Xiang, et al.
Published: (2026)
by: Hu, Yi-Xiang, et al.
Published: (2026)
Metric Criticality Identification for Cloud Microservices
by: Singal, Akanksha, et al.
Published: (2025)
by: Singal, Akanksha, et al.
Published: (2025)
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems
by: Zhou, Wenqing, et al.
Published: (2025)
by: Zhou, Wenqing, et al.
Published: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
by: Saurez, Enrique, et al.
Published: (2024)
by: Saurez, Enrique, et al.
Published: (2024)
Joint Resource Optimization, Computation Offloading and Resource Slicing for Multi-Edge Traffic-Cognitive Networks
by: Xiaoyang, Ting, et al.
Published: (2024)
by: Xiaoyang, Ting, et al.
Published: (2024)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
by: Chen, Daoyuan, et al.
Published: (2024)
by: Chen, Daoyuan, et al.
Published: (2024)
Chat AI: A Seamless Slurm-Native Solution for HPC-Based Services
by: Doosthosseini, Ali, et al.
Published: (2024)
by: Doosthosseini, Ali, et al.
Published: (2024)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
by: Mao, Ying, et al.
Published: (2020)
by: Mao, Ying, et al.
Published: (2020)
Comparison of Microservice Call Rate Predictions for Replication in the Cloud
by: Mehran, Narges, et al.
Published: (2023)
by: Mehran, Narges, et al.
Published: (2023)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
by: Samuthrsindh, Paramuth, et al.
Published: (2026)
by: Samuthrsindh, Paramuth, et al.
Published: (2026)
Adaptive Fault Tolerance Mechanisms of Large Language Models in Cloud Computing Environments
by: Jin, Yihong, et al.
Published: (2025)
by: Jin, Yihong, et al.
Published: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
by: Chen, Fahao, et al.
Published: (2025)
by: Chen, Fahao, et al.
Published: (2025)
MLCommons Cloud Masking Benchmark with Early Stopping
by: Chennamsetti, Varshitha, et al.
Published: (2023)
by: Chennamsetti, Varshitha, et al.
Published: (2023)
AI Factories: It's time to rethink the Cloud-HPC divide
by: Lopez, Pedro Garcia, et al.
Published: (2025)
by: Lopez, Pedro Garcia, et al.
Published: (2025)
The AI_INFN Platform: Artificial Intelligence Development in the Cloud
by: Anderlini, Lucio, et al.
Published: (2025)
by: Anderlini, Lucio, et al.
Published: (2025)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
A Meta-Heuristic Load Balancer for Cloud Computing Systems
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
Dynamic Resource Allocation for Virtual Machine Migration Optimization using Machine Learning
by: Gong, Yulu, et al.
Published: (2024)
by: Gong, Yulu, et al.
Published: (2024)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
by: Yang, Xinjun, et al.
Published: (2025)
by: Yang, Xinjun, et al.
Published: (2025)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
by: Qazi, Muhammad Azlan, et al.
Published: (2026)
by: Qazi, Muhammad Azlan, et al.
Published: (2026)
Towards using Reinforcement Learning for Scaling and Data Replication in Cloud Systems
by: Mokadem, Riad, et al.
Published: (2024)
by: Mokadem, Riad, et al.
Published: (2024)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
by: Xie, Jincheng, et al.
Published: (2026)
by: Xie, Jincheng, et al.
Published: (2026)
Optimal Configuration of API Resources in Cloud Native Computing
by: Truyen, Eddy, et al.
Published: (2025)
by: Truyen, Eddy, et al.
Published: (2025)
Reinforcement Learning-driven Data-intensive Workflow Scheduling for Volunteer Edge-Cloud
by: Mounesan, Motahare, et al.
Published: (2024)
by: Mounesan, Motahare, et al.
Published: (2024)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Similar Items
-
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
by: Liang, Mingyu, et al.
Published: (2025) -
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
by: Punniyamoorthy, Vinoth, et al.
Published: (2025) -
Adaptive AI-based Decentralized Resource Management in the Cloud-Edge Continuum
by: Li, Lanpei, et al.
Published: (2025) -
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025) -
Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
by: Legler, Julian, et al.
Published: (2025)