PET: An Annotated Dataset for Process Extraction from Natural Language Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bellan, Patrizio, van der Aa, Han, Dragoni, Mauro, Ghidini, Chiara, Ponzetto, Simone Paolo
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912493776928768
author Bellan, Patrizio
van der Aa, Han
Dragoni, Mauro
Ghidini, Chiara
Ponzetto, Simone Paolo
author_facet Bellan, Patrizio
van der Aa, Han
Dragoni, Mauro
Ghidini, Chiara
Ponzetto, Simone Paolo
contents Process extraction from text is an important task of process discovery, for which various approaches have been developed in recent years. However, in contrast to other information extraction tasks, there is a lack of gold-standard corpora of business process descriptions that are carefully annotated with all the entities and relationships of interest. Due to this, it is currently hard to compare the results obtained by extraction approaches in an objective manner, whereas the lack of annotated texts also prevents the application of data-driven information extraction methodologies, typical of the natural language processing field. Therefore, to bridge this gap, we present the PET dataset, a first corpus of business process descriptions annotated with activities, gateways, actors, and flow information. We present our new resource, including a variety of baselines to benchmark the difficulty and challenges of business process extraction from text. PET can be accessed via huggingface.co/datasets/patriziobellan/PET
format Preprint
id arxiv_https___arxiv_org_abs_2203_04860
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle PET: An Annotated Dataset for Process Extraction from Natural Language Text
Bellan, Patrizio
van der Aa, Han
Dragoni, Mauro
Ghidini, Chiara
Ponzetto, Simone Paolo
Computation and Language
Process extraction from text is an important task of process discovery, for which various approaches have been developed in recent years. However, in contrast to other information extraction tasks, there is a lack of gold-standard corpora of business process descriptions that are carefully annotated with all the entities and relationships of interest. Due to this, it is currently hard to compare the results obtained by extraction approaches in an objective manner, whereas the lack of annotated texts also prevents the application of data-driven information extraction methodologies, typical of the natural language processing field. Therefore, to bridge this gap, we present the PET dataset, a first corpus of business process descriptions annotated with activities, gateways, actors, and flow information. We present our new resource, including a variety of baselines to benchmark the difficulty and challenges of business process extraction from text. PET can be accessed via huggingface.co/datasets/patriziobellan/PET
title PET: An Annotated Dataset for Process Extraction from Natural Language Text
topic Computation and Language
url https://arxiv.org/abs/2203.04860