N-Version Assessment and Enhancement of Generative AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kessel, Marcus, Atkinson, Colin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909357857308672
author Kessel, Marcus
Atkinson, Colin
author_facet Kessel, Marcus
Atkinson, Colin
contents Generative AI (GAI) holds great potential to improve software engineering productivity, but its untrustworthy outputs, particularly in code synthesis, pose significant challenges. The need for extensive verification and validation (V&V) of GAI-generated artifacts may undermine the potential productivity gains. This paper proposes a way of mitigating these risks by exploiting GAI's ability to generate multiple versions of code and tests to facilitate comparative analysis across versions. Rather than relying on the quality of a single test or code module, this "differential GAI" (D-GAI) approach promotes more reliable quality evaluation through version diversity. We introduce the Large-Scale Software Observatorium (LASSO), a platform that supports D-GAI by executing and analyzing large sets of code versions and tests. We discuss how LASSO enables rigorous evaluation of GAI-generated artifacts and propose its application in both software development and GAI research.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle N-Version Assessment and Enhancement of Generative AI
Kessel, Marcus
Atkinson, Colin
Software Engineering
Artificial Intelligence
D.2.1; D.2.4; I.2.2; I.2.7
Generative AI (GAI) holds great potential to improve software engineering productivity, but its untrustworthy outputs, particularly in code synthesis, pose significant challenges. The need for extensive verification and validation (V&V) of GAI-generated artifacts may undermine the potential productivity gains. This paper proposes a way of mitigating these risks by exploiting GAI's ability to generate multiple versions of code and tests to facilitate comparative analysis across versions. Rather than relying on the quality of a single test or code module, this "differential GAI" (D-GAI) approach promotes more reliable quality evaluation through version diversity. We introduce the Large-Scale Software Observatorium (LASSO), a platform that supports D-GAI by executing and analyzing large sets of code versions and tests. We discuss how LASSO enables rigorous evaluation of GAI-generated artifacts and propose its application in both software development and GAI research.
title N-Version Assessment and Enhancement of Generative AI
topic Software Engineering
Artificial Intelligence
D.2.1; D.2.4; I.2.2; I.2.7
url https://arxiv.org/abs/2409.14071