A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bastos, Anson, Venneti, Shreeya, Parayil, Anjaly, Choure, Ayush, Bansal, Chetan, Wang, Rujia
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911509417820160
author Bastos, Anson
Venneti, Shreeya
Parayil, Anjaly
Choure, Ayush
Bansal, Chetan
Wang, Rujia
author_facet Bastos, Anson
Venneti, Shreeya
Parayil, Anjaly
Choure, Ayush
Bansal, Chetan
Wang, Rujia
contents Reliability of large-scale cloud services is critical for user satisfaction and business continuity. Despite significant investments in reliability engineering, production incidents remain inevitable, often leading to customer impact and operational overhead. In large cloud companies, multiple services are deployed across regions necessitating robust health monitoring systems. However, the current monitor configuration process is manual, largely reactive and ad hoc, resulting in gaps in coverage and redundant alerts. In this paper, we present a comprehensive study of monitor creation in Microsoft, identifying key components in the existing process. We further design a modular recommendation framework that processes the graph structured service entities to suggest optimal monitor configurations. Through extensive experimentation on historical data and user study of recommendations for production services at Microsoft, we demonstrate the efficacy of our approach in providing relevant recommendations for monitor configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12268
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
Bastos, Anson
Venneti, Shreeya
Parayil, Anjaly
Choure, Ayush
Bansal, Chetan
Wang, Rujia
Distributed, Parallel, and Cluster Computing
Machine Learning
Reliability of large-scale cloud services is critical for user satisfaction and business continuity. Despite significant investments in reliability engineering, production incidents remain inevitable, often leading to customer impact and operational overhead. In large cloud companies, multiple services are deployed across regions necessitating robust health monitoring systems. However, the current monitor configuration process is manual, largely reactive and ad hoc, resulting in gaps in coverage and redundant alerts. In this paper, we present a comprehensive study of monitor creation in Microsoft, identifying key components in the existing process. We further design a modular recommendation framework that processes the graph structured service entities to suggest optimal monitor configurations. Through extensive experimentation on historical data and user study of recommendations for production services at Microsoft, we demonstrate the efficacy of our approach in providing relevant recommendations for monitor configurations.
title A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2603.12268