Applying Process Mining on Scientific Workflows: a Case Study on High Performance Computing Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sadeghibogar, Zahra, Berti, Alessandro, Pegoraro, Marco, van der Aalst, Wil M. P.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913689949437952
author Sadeghibogar, Zahra
Berti, Alessandro
Pegoraro, Marco
van der Aalst, Wil M. P.
author_facet Sadeghibogar, Zahra
Berti, Alessandro
Pegoraro, Marco
van der Aalst, Wil M. P.
contents Computer-based scientific experiments are becoming increasingly data-intensive, necessitating the use of High-Performance Computing (HPC) clusters to handle large scientific workflows. These workflows result in complex data and control flows within the system, making analysis challenging. This paper focuses on the extraction of case IDs from SLURM-based HPC cluster logs, a crucial step for applying mainstream process mining techniques. The core contribution is the development of methods to correlate jobs in the system, whether their interdependencies are explicitly specified or not. We present our log extraction and correlation techniques, supported by experiments that validate our approach, enabling comprehensive documentation of workflows and identification of performance bottlenecks.
format Preprint
id arxiv_https___arxiv_org_abs_2307_02833
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Applying Process Mining on Scientific Workflows: a Case Study on High Performance Computing Data
Sadeghibogar, Zahra
Berti, Alessandro
Pegoraro, Marco
van der Aalst, Wil M. P.
Databases
Computer-based scientific experiments are becoming increasingly data-intensive, necessitating the use of High-Performance Computing (HPC) clusters to handle large scientific workflows. These workflows result in complex data and control flows within the system, making analysis challenging. This paper focuses on the extraction of case IDs from SLURM-based HPC cluster logs, a crucial step for applying mainstream process mining techniques. The core contribution is the development of methods to correlate jobs in the system, whether their interdependencies are explicitly specified or not. We present our log extraction and correlation techniques, supported by experiments that validate our approach, enabling comprehensive documentation of workflows and identification of performance bottlenecks.
title Applying Process Mining on Scientific Workflows: a Case Study on High Performance Computing Data
topic Databases
url https://arxiv.org/abs/2307.02833