Saved in:
Bibliographic Details
Main Authors: Yao, Guanzi, Liu, Heyao, Dai, Linyan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.10253
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909736650145792
author Yao, Guanzi
Liu, Heyao
Dai, Linyan
author_facet Yao, Guanzi
Liu, Heyao
Dai, Linyan
contents This paper addresses the challenges of high resource dynamism and scheduling complexity in cloud-native database systems. It proposes an adaptive resource orchestration method based on multi-agent reinforcement learning. The method introduces a heterogeneous role-based agent modeling mechanism. This allows different resource entities, such as compute nodes, storage nodes, and schedulers, to adopt distinct policy representations. These agents are better able to reflect diverse functional responsibilities and local environmental characteristics within the system. A reward-shaping mechanism is designed to integrate local observations with global feedback. This helps mitigate policy learning bias caused by incomplete state observations. By combining real-time local performance signals with global system value estimation, the mechanism improves coordination among agents and enhances policy convergence stability. A unified multi-agent training framework is developed and evaluated on a representative production scheduling dataset. Experimental results show that the proposed method outperforms traditional approaches across multiple key metrics. These include resource utilization, scheduling latency, policy convergence speed, system stability, and fairness. The results demonstrate strong generalization and practical utility. Across various experimental scenarios, the method proves effective in handling orchestration tasks with high concurrency, high-dimensional state spaces, and complex dependency relationships. This confirms its advantages in real-world, large-scale scheduling environments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10253
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Agent Reinforcement Learning for Adaptive Resource Orchestration in Cloud-Native Clusters
Yao, Guanzi
Liu, Heyao
Dai, Linyan
Machine Learning
This paper addresses the challenges of high resource dynamism and scheduling complexity in cloud-native database systems. It proposes an adaptive resource orchestration method based on multi-agent reinforcement learning. The method introduces a heterogeneous role-based agent modeling mechanism. This allows different resource entities, such as compute nodes, storage nodes, and schedulers, to adopt distinct policy representations. These agents are better able to reflect diverse functional responsibilities and local environmental characteristics within the system. A reward-shaping mechanism is designed to integrate local observations with global feedback. This helps mitigate policy learning bias caused by incomplete state observations. By combining real-time local performance signals with global system value estimation, the mechanism improves coordination among agents and enhances policy convergence stability. A unified multi-agent training framework is developed and evaluated on a representative production scheduling dataset. Experimental results show that the proposed method outperforms traditional approaches across multiple key metrics. These include resource utilization, scheduling latency, policy convergence speed, system stability, and fairness. The results demonstrate strong generalization and practical utility. Across various experimental scenarios, the method proves effective in handling orchestration tasks with high concurrency, high-dimensional state spaces, and complex dependency relationships. This confirms its advantages in real-world, large-scale scheduling environments.
title Multi-Agent Reinforcement Learning for Adaptive Resource Orchestration in Cloud-Native Clusters
topic Machine Learning
url https://arxiv.org/abs/2508.10253