Systematic approach to root cause analysis in distributed data processing systems

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Bollineni, Satyadeepak
Format: Recurso digital
Language:English
Published: Zenodo 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866902308527276032
author Bollineni, Satyadeepak
author_facet Bollineni, Satyadeepak
contents <p>Distributed data processing is a powerful capability, but with it comes the challenge of ensuring the reliability and performance of the system often on a larger scale, it is especially important to systematically identify the root cause of failures and address them accordingly. Cloud computing has changed the game by introducing scale, flexibility and low-cost alternatives to big data processing. With distributed systems getting increasingly complex, diagnosing failures has become defeated due to many components relying on each other and as workloads change dynamically. This paper presents a systematic approach for performing root cause analysis (RCA) in a distributed setting one that covers automatic monitoring, anomaly detection, and log-based analytics. Overcoming the RCA challenges with cloud-native tools like Azure Data Factory, Power BI, and anomaly detection through machine learning are discussed. The research also discusses best practices for reducing downtime and performance optimization with predictive maintenance strategy. Cloud technologies have enabled organizations to achieve greater operational efficiency through better system resilience and decision-making in modern data-driven environment.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17112578
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Systematic approach to root cause analysis in distributed data processing systems
Bollineni, Satyadeepak
Root Cause Analysis
Distributed Data Processing
Cloud Computing
Anomaly Detection
Predictive Maintenance
Azure Data Factory
Power Bi
System Resilience
Log Analytics
<p>Distributed data processing is a powerful capability, but with it comes the challenge of ensuring the reliability and performance of the system often on a larger scale, it is especially important to systematically identify the root cause of failures and address them accordingly. Cloud computing has changed the game by introducing scale, flexibility and low-cost alternatives to big data processing. With distributed systems getting increasingly complex, diagnosing failures has become defeated due to many components relying on each other and as workloads change dynamically. This paper presents a systematic approach for performing root cause analysis (RCA) in a distributed setting one that covers automatic monitoring, anomaly detection, and log-based analytics. Overcoming the RCA challenges with cloud-native tools like Azure Data Factory, Power BI, and anomaly detection through machine learning are discussed. The research also discusses best practices for reducing downtime and performance optimization with predictive maintenance strategy. Cloud technologies have enabled organizations to achieve greater operational efficiency through better system resilience and decision-making in modern data-driven environment.</p>
title Systematic approach to root cause analysis in distributed data processing systems
topic Root Cause Analysis
Distributed Data Processing
Cloud Computing
Anomaly Detection
Predictive Maintenance
Azure Data Factory
Power Bi
System Resilience
Log Analytics
url https://doi.org/10.5281/zenodo.17112578