LLMs for Science: Usage for Code Generation and Data Analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nejjar, Mohamed, Zacharias, Luca, Stiehle, Fabian, Weber, Ingo
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914766260273152
author Nejjar, Mohamed
Zacharias, Luca
Stiehle, Fabian
Weber, Ingo
author_facet Nejjar, Mohamed
Zacharias, Luca
Stiehle, Fabian
Weber, Ingo
contents Large language models (LLMs) have been touted to enable increased productivity in many areas of today's work life. Scientific research as an area of work is no exception: the potential of LLM-based tools to assist in the daily work of scientists has become a highly discussed topic across disciplines. However, we are only at the very onset of this subject of study. It is still unclear how the potential of LLMs will materialise in research practice. With this study, we give first empirical evidence on the use of LLMs in the research process. We have investigated a set of use cases for LLM-based tools in scientific research, and conducted a first study to assess to which degree current tools are helpful. In this paper we report specifically on use cases related to software engineering, such as generating application code and developing scripts for data analytics. While we studied seemingly simple use cases, results across tools differ significantly. Our results highlight the promise of LLM-based tools in general, yet we also observe various issues, particularly regarding the integrity of the output these tools provide.
format Preprint
id arxiv_https___arxiv_org_abs_2311_16733
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LLMs for Science: Usage for Code Generation and Data Analysis
Nejjar, Mohamed
Zacharias, Luca
Stiehle, Fabian
Weber, Ingo
Software Engineering
Artificial Intelligence
Computation and Language
Large language models (LLMs) have been touted to enable increased productivity in many areas of today's work life. Scientific research as an area of work is no exception: the potential of LLM-based tools to assist in the daily work of scientists has become a highly discussed topic across disciplines. However, we are only at the very onset of this subject of study. It is still unclear how the potential of LLMs will materialise in research practice. With this study, we give first empirical evidence on the use of LLMs in the research process. We have investigated a set of use cases for LLM-based tools in scientific research, and conducted a first study to assess to which degree current tools are helpful. In this paper we report specifically on use cases related to software engineering, such as generating application code and developing scripts for data analytics. While we studied seemingly simple use cases, results across tools differ significantly. Our results highlight the promise of LLM-based tools in general, yet we also observe various issues, particularly regarding the integrity of the output these tools provide.
title LLMs for Science: Usage for Code Generation and Data Analysis
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2311.16733