Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Diehl, Patrick, Nader, Noujoud, Brandt, Steve, Kaiser, Hartmut
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908356665409536
author Diehl, Patrick
Nader, Noujoud
Brandt, Steve
Kaiser, Hartmut
author_facet Diehl, Patrick
Nader, Noujoud
Brandt, Steve
Kaiser, Hartmut
contents This study evaluates the capabilities of ChatGPT versions 3.5 and 4 in generating code across a diverse range of programming languages. Our objective is to assess the effectiveness of these AI models for generating scientific programs. To this end, we asked ChatGPT to generate three distinct codes: a simple numerical integration, a conjugate gradient solver, and a parallel 1D stencil-based heat equation solver. The focus of our analysis was on the compilation, runtime performance, and accuracy of the codes. While both versions of ChatGPT successfully created codes that compiled and ran (with some help), some languages were easier for the AI to use than others (possibly because of the size of the training sets used). Parallel codes -- even the simple example we chose to study here -- also difficult for the AI to generate correctly.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13101
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
Diehl, Patrick
Nader, Noujoud
Brandt, Steve
Kaiser, Hartmut
Software Engineering
Artificial Intelligence
This study evaluates the capabilities of ChatGPT versions 3.5 and 4 in generating code across a diverse range of programming languages. Our objective is to assess the effectiveness of these AI models for generating scientific programs. To this end, we asked ChatGPT to generate three distinct codes: a simple numerical integration, a conjugate gradient solver, and a parallel 1D stencil-based heat equation solver. The focus of our analysis was on the compilation, runtime performance, and accuracy of the codes. While both versions of ChatGPT successfully created codes that compiled and ran (with some help), some languages were easier for the AI to use than others (possibly because of the size of the training sets used). Parallel codes -- even the simple example we chose to study here -- also difficult for the AI to generate correctly.
title Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2405.13101