Leveraging Generative AI for Extracting Process Models from Multimodal Documents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910476670074880 |
|---|---|
| author | Voelter, Marvin Hadian, Raheleh Kampik, Timotheus Breitmayer, Marius Reichert, Manfred |
| author_facet | Voelter, Marvin Hadian, Raheleh Kampik, Timotheus Breitmayer, Marius Reichert, Manfred |
| contents | This paper presents an investigation of the capabilities of Generative Pre-trained Transformers (GPTs) to auto-generate graphical process models from multi-modal (i.e., text- and image-based) inputs. More precisely, we first introduce a small dataset as well as a set of evaluation metrics that allow for a ground truth-based evaluation of multi-modal process model generation capabilities. We then conduct an initial evaluation of commercial GPT capabilities using zero-, one-, and few-shot prompting strategies. Our results indicate that GPTs can be useful tools for semi-automated process modeling based on multi-modal inputs. More importantly, the dataset and evaluation metrics as well as the open-source evaluation code provide a structured framework for continued systematic evaluations moving forward. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_04959 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Leveraging Generative AI for Extracting Process Models from Multimodal Documents Voelter, Marvin Hadian, Raheleh Kampik, Timotheus Breitmayer, Marius Reichert, Manfred Software Engineering This paper presents an investigation of the capabilities of Generative Pre-trained Transformers (GPTs) to auto-generate graphical process models from multi-modal (i.e., text- and image-based) inputs. More precisely, we first introduce a small dataset as well as a set of evaluation metrics that allow for a ground truth-based evaluation of multi-modal process model generation capabilities. We then conduct an initial evaluation of commercial GPT capabilities using zero-, one-, and few-shot prompting strategies. Our results indicate that GPTs can be useful tools for semi-automated process modeling based on multi-modal inputs. More importantly, the dataset and evaluation metrics as well as the open-source evaluation code provide a structured framework for continued systematic evaluations moving forward. |
| title | Leveraging Generative AI for Extracting Process Models from Multimodal Documents |
| topic | Software Engineering |
| url | https://arxiv.org/abs/2406.04959 |