Tailoring Machine Learning for Process Mining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ceravolo, Paolo, Junior, Sylvio Barbon, Damiani, Ernesto, van der Aalst, Wil
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913237333704704
author Ceravolo, Paolo
Junior, Sylvio Barbon
Damiani, Ernesto
van der Aalst, Wil
author_facet Ceravolo, Paolo
Junior, Sylvio Barbon
Damiani, Ernesto
van der Aalst, Wil
contents Machine learning models are routinely integrated into process mining pipelines to carry out tasks like data transformation, noise reduction, anomaly detection, classification, and prediction. Often, the design of such models is based on some ad-hoc assumptions about the corresponding data distributions, which are not necessarily in accordance with the non-parametric distributions typically observed with process data. Moreover, the learning procedure they follow ignores the constraints concurrency imposes to process data. Data encoding is a key element to smooth the mismatch between these assumptions but its potential is poorly exploited. In this paper, we argue that a deeper insight into the issues raised by training machine learning models with process data is crucial to ground a sound integration of process mining and machine learning. Our analysis of such issues is aimed at laying the foundation for a methodology aimed at correctly aligning machine learning with process mining requirements and stimulating the research to elaborate in this direction.
format Preprint
id arxiv_https___arxiv_org_abs_2306_10341
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Tailoring Machine Learning for Process Mining
Ceravolo, Paolo
Junior, Sylvio Barbon
Damiani, Ernesto
van der Aalst, Wil
Machine Learning
Artificial Intelligence
Databases
68
I.2.6
Machine learning models are routinely integrated into process mining pipelines to carry out tasks like data transformation, noise reduction, anomaly detection, classification, and prediction. Often, the design of such models is based on some ad-hoc assumptions about the corresponding data distributions, which are not necessarily in accordance with the non-parametric distributions typically observed with process data. Moreover, the learning procedure they follow ignores the constraints concurrency imposes to process data. Data encoding is a key element to smooth the mismatch between these assumptions but its potential is poorly exploited. In this paper, we argue that a deeper insight into the issues raised by training machine learning models with process data is crucial to ground a sound integration of process mining and machine learning. Our analysis of such issues is aimed at laying the foundation for a methodology aimed at correctly aligning machine learning with process mining requirements and stimulating the research to elaborate in this direction.
title Tailoring Machine Learning for Process Mining
topic Machine Learning
Artificial Intelligence
Databases
68
I.2.6
url https://arxiv.org/abs/2306.10341