Automated BPMN Model Generation from Textual Process Descriptions: A Multi-Stage LLM-Driven Approach

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Matei, Ion, Zhenirovskyy, Maksym, Sekar, Praveen Kumar Menaka, Wong, Hon Yung
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910127570812928
author Matei, Ion
Zhenirovskyy, Maksym
Sekar, Praveen Kumar Menaka
Wong, Hon Yung
author_facet Matei, Ion
Zhenirovskyy, Maksym
Sekar, Praveen Kumar Menaka
Wong, Hon Yung
contents Automatically reconstructing BPMN models from unstructured natural-language descriptions remains challenging due to heterogeneous modeling conventions, multilingual sources, and the lack of reliable ground truth. We present a scalable, multi-stage LLM-driven pipeline that automates both ground-truth construction and model reconstruction. Multilingual BPMN XML files are translated into English, validated using execution-oriented compliance checks in SpiffWorkflow, and iteratively repaired through targeted LLM-guided corrections to produce a consistent ground-truth corpus. From these validated models, process descriptions are generated and used to reconstruct executable BPMN~2.0 XML diagrams without manual curation. We introduce a multi-dimensional similarity framework combining structural metrics, type-distribution alignment, and embedding-based semantic measures. In an empirical study of 750 public BPMN diagrams, the pipeline generated 387 validated ground-truth models and achieved average reconstruction similarity above 0.75, including approximately 50 near-perfect reconstructions differing only in minor naming variations. The results demonstrate that LLMs can generate structurally compliant and semantically meaningful BPMN diagrams at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12105
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Automated BPMN Model Generation from Textual Process Descriptions: A Multi-Stage LLM-Driven Approach
Matei, Ion
Zhenirovskyy, Maksym
Sekar, Praveen Kumar Menaka
Wong, Hon Yung
Software Engineering
Automatically reconstructing BPMN models from unstructured natural-language descriptions remains challenging due to heterogeneous modeling conventions, multilingual sources, and the lack of reliable ground truth. We present a scalable, multi-stage LLM-driven pipeline that automates both ground-truth construction and model reconstruction. Multilingual BPMN XML files are translated into English, validated using execution-oriented compliance checks in SpiffWorkflow, and iteratively repaired through targeted LLM-guided corrections to produce a consistent ground-truth corpus. From these validated models, process descriptions are generated and used to reconstruct executable BPMN~2.0 XML diagrams without manual curation. We introduce a multi-dimensional similarity framework combining structural metrics, type-distribution alignment, and embedding-based semantic measures. In an empirical study of 750 public BPMN diagrams, the pipeline generated 387 validated ground-truth models and achieved average reconstruction similarity above 0.75, including approximately 50 near-perfect reconstructions differing only in minor naming variations. The results demonstrate that LLMs can generate structurally compliant and semantically meaningful BPMN diagrams at scale.
title Automated BPMN Model Generation from Textual Process Descriptions: A Multi-Stage LLM-Driven Approach
topic Software Engineering
url https://arxiv.org/abs/2604.12105