Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911123550240768 |
|---|---|
| author | Islam, Mohammad Saiful Rakha, Mohamed Sami Pourmajidi, William Sivaloganathan, Janakan Steinbacher, John Miranskyy, Andriy |
| author_facet | Islam, Mohammad Saiful Rakha, Mohamed Sami Pourmajidi, William Sivaloganathan, Janakan Steinbacher, John Miranskyy, Andriy |
| contents | As Large-Scale Cloud Systems (LCS) become increasingly complex, effective anomaly detection is critical for ensuring system reliability and performance. However, there is a shortage of large-scale, real-world datasets available for benchmarking anomaly detection methods.
To address this gap, we introduce a new high-dimensional dataset from IBM Cloud, collected over 4.5 months from the IBM Cloud Console. This dataset comprises 39,365 rows and 117,448 columns of telemetry data. Additionally, we demonstrate the application of machine learning models for anomaly detection and discuss the key challenges faced in this process.
This study and the accompanying dataset provide a resource for researchers and practitioners in cloud system monitoring. It facilitates more efficient testing of anomaly detection methods in real-world data, helping to advance the development of robust solutions to maintain the health and performance of large-scale cloud infrastructures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_09047 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset Islam, Mohammad Saiful Rakha, Mohamed Sami Pourmajidi, William Sivaloganathan, Janakan Steinbacher, John Miranskyy, Andriy Machine Learning Distributed, Parallel, and Cluster Computing Software Engineering As Large-Scale Cloud Systems (LCS) become increasingly complex, effective anomaly detection is critical for ensuring system reliability and performance. However, there is a shortage of large-scale, real-world datasets available for benchmarking anomaly detection methods. To address this gap, we introduce a new high-dimensional dataset from IBM Cloud, collected over 4.5 months from the IBM Cloud Console. This dataset comprises 39,365 rows and 117,448 columns of telemetry data. Additionally, we demonstrate the application of machine learning models for anomaly detection and discuss the key challenges faced in this process. This study and the accompanying dataset provide a resource for researchers and practitioners in cloud system monitoring. It facilitates more efficient testing of anomaly detection methods in real-world data, helping to advance the development of robust solutions to maintain the health and performance of large-scale cloud infrastructures. |
| title | Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing Software Engineering |
| url | https://arxiv.org/abs/2411.09047 |