Let's Simply Count: Quantifying Distributional Similarity Between Activities in Event Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kirchmann, Henrik, Fahrenkrog-Petersen, Stephan A., Lu, Xixi, Weidlich, Matthias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914032572694528
author Kirchmann, Henrik
Fahrenkrog-Petersen, Stephan A.
Lu, Xixi
Weidlich, Matthias
author_facet Kirchmann, Henrik
Fahrenkrog-Petersen, Stephan A.
Lu, Xixi
Weidlich, Matthias
contents To obtain insights from event data, advanced process mining methods assess the similarity of activities to incorporate their semantic relations into the analysis. Here, distributional similarity that captures similarity from activity co-occurrences is commonly employed. However, existing work for distributional similarity in process mining adopt neural network-based approaches as developed for natural language processing, e.g., word2vec and autoencoders. While these approaches have been shown to be effective, their downsides are high computational costs and limited interpretability of the learned representations. In this work, we argue for simplicity in the modeling of distributional similarity of activities. We introduce count-based embeddings that avoid a complex training process and offer a direct interpretable representation. To underpin our call for simple embeddings, we contribute a comprehensive benchmarking framework, which includes means to assess the intrinsic quality of embeddings, their performance in downstream applications, and their computational efficiency. In experiments that compare against the state of the art, we demonstrate that count-based embeddings provide a highly effective and efficient basis for distributional similarity between activities in event data.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09440
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let's Simply Count: Quantifying Distributional Similarity Between Activities in Event Data
Kirchmann, Henrik
Fahrenkrog-Petersen, Stephan A.
Lu, Xixi
Weidlich, Matthias
Databases
To obtain insights from event data, advanced process mining methods assess the similarity of activities to incorporate their semantic relations into the analysis. Here, distributional similarity that captures similarity from activity co-occurrences is commonly employed. However, existing work for distributional similarity in process mining adopt neural network-based approaches as developed for natural language processing, e.g., word2vec and autoencoders. While these approaches have been shown to be effective, their downsides are high computational costs and limited interpretability of the learned representations. In this work, we argue for simplicity in the modeling of distributional similarity of activities. We introduce count-based embeddings that avoid a complex training process and offer a direct interpretable representation. To underpin our call for simple embeddings, we contribute a comprehensive benchmarking framework, which includes means to assess the intrinsic quality of embeddings, their performance in downstream applications, and their computational efficiency. In experiments that compare against the state of the art, we demonstrate that count-based embeddings provide a highly effective and efficient basis for distributional similarity between activities in event data.
title Let's Simply Count: Quantifying Distributional Similarity Between Activities in Event Data
topic Databases
url https://arxiv.org/abs/2509.09440