Evaluating Large Language Models in Process Mining: Capabilities, Benchmarks, and Evaluation Strategies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berti, Alessandro, Kourani, Humam, Hafke, Hannes, Li, Chiao-Yun, Schuster, Daniel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911935358828544
author Berti, Alessandro
Kourani, Humam
Hafke, Hannes
Li, Chiao-Yun
Schuster, Daniel
author_facet Berti, Alessandro
Kourani, Humam
Hafke, Hannes
Li, Chiao-Yun
Schuster, Daniel
contents Using Large Language Models (LLMs) for Process Mining (PM) tasks is becoming increasingly essential, and initial approaches yield promising results. However, little attention has been given to developing strategies for evaluating and benchmarking the utility of incorporating LLMs into PM tasks. This paper reviews the current implementations of LLMs in PM and reflects on three different questions. 1) What is the minimal set of capabilities required for PM on LLMs? 2) Which benchmark strategies help choose optimal LLMs for PM? 3) How do we evaluate the output of LLMs on specific PM tasks? The answer to these questions is fundamental to the development of comprehensive process mining benchmarks on LLMs covering different tasks and implementation paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06749
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Large Language Models in Process Mining: Capabilities, Benchmarks, and Evaluation Strategies
Berti, Alessandro
Kourani, Humam
Hafke, Hannes
Li, Chiao-Yun
Schuster, Daniel
Databases
Using Large Language Models (LLMs) for Process Mining (PM) tasks is becoming increasingly essential, and initial approaches yield promising results. However, little attention has been given to developing strategies for evaluating and benchmarking the utility of incorporating LLMs into PM tasks. This paper reviews the current implementations of LLMs in PM and reflects on three different questions. 1) What is the minimal set of capabilities required for PM on LLMs? 2) Which benchmark strategies help choose optimal LLMs for PM? 3) How do we evaluate the output of LLMs on specific PM tasks? The answer to these questions is fundamental to the development of comprehensive process mining benchmarks on LLMs covering different tasks and implementation paradigms.
title Evaluating Large Language Models in Process Mining: Capabilities, Benchmarks, and Evaluation Strategies
topic Databases
url https://arxiv.org/abs/2403.06749