Leveraging Generative AI for Extracting Process Models from Multimodal Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Voelter, Marvin, Hadian, Raheleh, Kampik, Timotheus, Breitmayer, Marius, Reichert, Manfred
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910476670074880
author Voelter, Marvin
Hadian, Raheleh
Kampik, Timotheus
Breitmayer, Marius
Reichert, Manfred
author_facet Voelter, Marvin
Hadian, Raheleh
Kampik, Timotheus
Breitmayer, Marius
Reichert, Manfred
contents This paper presents an investigation of the capabilities of Generative Pre-trained Transformers (GPTs) to auto-generate graphical process models from multi-modal (i.e., text- and image-based) inputs. More precisely, we first introduce a small dataset as well as a set of evaluation metrics that allow for a ground truth-based evaluation of multi-modal process model generation capabilities. We then conduct an initial evaluation of commercial GPT capabilities using zero-, one-, and few-shot prompting strategies. Our results indicate that GPTs can be useful tools for semi-automated process modeling based on multi-modal inputs. More importantly, the dataset and evaluation metrics as well as the open-source evaluation code provide a structured framework for continued systematic evaluations moving forward.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04959
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Generative AI for Extracting Process Models from Multimodal Documents
Voelter, Marvin
Hadian, Raheleh
Kampik, Timotheus
Breitmayer, Marius
Reichert, Manfred
Software Engineering
This paper presents an investigation of the capabilities of Generative Pre-trained Transformers (GPTs) to auto-generate graphical process models from multi-modal (i.e., text- and image-based) inputs. More precisely, we first introduce a small dataset as well as a set of evaluation metrics that allow for a ground truth-based evaluation of multi-modal process model generation capabilities. We then conduct an initial evaluation of commercial GPT capabilities using zero-, one-, and few-shot prompting strategies. Our results indicate that GPTs can be useful tools for semi-automated process modeling based on multi-modal inputs. More importantly, the dataset and evaluation metrics as well as the open-source evaluation code provide a structured framework for continued systematic evaluations moving forward.
title Leveraging Generative AI for Extracting Process Models from Multimodal Documents
topic Software Engineering
url https://arxiv.org/abs/2406.04959