Large Language Models to the Rescue: Reducing the Complexity in Scientific Workflow Development Using ChatGPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sänger, Mario, De Mecquenem, Ninon, Lewińska, Katarzyna Ewa, Bountris, Vasilis, Lehmann, Fabian, Leser, Ulf, Kosch, Thomas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909244883730432
author Sänger, Mario
De Mecquenem, Ninon
Lewińska, Katarzyna Ewa
Bountris, Vasilis
Lehmann, Fabian
Leser, Ulf
Kosch, Thomas
author_facet Sänger, Mario
De Mecquenem, Ninon
Lewińska, Katarzyna Ewa
Bountris, Vasilis
Lehmann, Fabian
Leser, Ulf
Kosch, Thomas
contents Scientific workflow systems are increasingly popular for expressing and executing complex data analysis pipelines over large datasets, as they offer reproducibility, dependability, and scalability of analyses by automatic parallelization on large compute clusters. However, implementing workflows is difficult due to the involvement of many black-box tools and the deep infrastructure stack necessary for their execution. Simultaneously, user-supporting tools are rare, and the number of available examples is much lower than in classical programming languages. To address these challenges, we investigate the efficiency of Large Language Models (LLMs), specifically ChatGPT, to support users when dealing with scientific workflows. We performed three user studies in two scientific domains to evaluate ChatGPT for comprehending, adapting, and extending workflows. Our results indicate that LLMs efficiently interpret workflows but achieve lower performance for exchanging components or purposeful workflow extensions. We characterize their limitations in these challenging scenarios and suggest future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2311_01825
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Large Language Models to the Rescue: Reducing the Complexity in Scientific Workflow Development Using ChatGPT
Sänger, Mario
De Mecquenem, Ninon
Lewińska, Katarzyna Ewa
Bountris, Vasilis
Lehmann, Fabian
Leser, Ulf
Kosch, Thomas
Distributed, Parallel, and Cluster Computing
Computation and Language
Human-Computer Interaction
Scientific workflow systems are increasingly popular for expressing and executing complex data analysis pipelines over large datasets, as they offer reproducibility, dependability, and scalability of analyses by automatic parallelization on large compute clusters. However, implementing workflows is difficult due to the involvement of many black-box tools and the deep infrastructure stack necessary for their execution. Simultaneously, user-supporting tools are rare, and the number of available examples is much lower than in classical programming languages. To address these challenges, we investigate the efficiency of Large Language Models (LLMs), specifically ChatGPT, to support users when dealing with scientific workflows. We performed three user studies in two scientific domains to evaluate ChatGPT for comprehending, adapting, and extending workflows. Our results indicate that LLMs efficiently interpret workflows but achieve lower performance for exchanging components or purposeful workflow extensions. We characterize their limitations in these challenging scenarios and suggest future research directions.
title Large Language Models to the Rescue: Reducing the Complexity in Scientific Workflow Development Using ChatGPT
topic Distributed, Parallel, and Cluster Computing
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2311.01825