Assured LLM-Based Software Engineering

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Alshahwan, Nadia, Harman, Mark, Harper, Inna, Marginean, Alexandru, Sengupta, Shubho, Wang, Eddy
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914669781843968
author Alshahwan, Nadia
Harman, Mark
Harper, Inna
Marginean, Alexandru
Sengupta, Shubho
Wang, Eddy
author_facet Alshahwan, Nadia
Harman, Mark
Harper, Inna
Marginean, Alexandru
Sengupta, Shubho
Wang, Eddy
contents In this paper we address the following question: How can we use Large Language Models (LLMs) to improve code independently of a human, while ensuring that the improved code - does not regress the properties of the original code? - improves the original in a verifiable and measurable way? To address this question, we advocate Assured LLM-Based Software Engineering; a generate-and-test approach, inspired by Genetic Improvement. Assured LLMSE applies a series of semantic filters that discard code that fails to meet these twin guarantees. This overcomes the potential problem of LLM's propensity to hallucinate. It allows us to generate code using LLMs, independently of any human. The human plays the role only of final code reviewer, as they would do with code generated by other human engineers. This paper is an outline of the content of the keynote by Mark Harman at the International Workshop on Interpretability, Robustness, and Benchmarking in Neural Software Engineering, Monday 15th April 2024, Lisbon, Portugal.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04380
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assured LLM-Based Software Engineering
Alshahwan, Nadia
Harman, Mark
Harper, Inna
Marginean, Alexandru
Sengupta, Shubho
Wang, Eddy
Software Engineering
In this paper we address the following question: How can we use Large Language Models (LLMs) to improve code independently of a human, while ensuring that the improved code - does not regress the properties of the original code? - improves the original in a verifiable and measurable way? To address this question, we advocate Assured LLM-Based Software Engineering; a generate-and-test approach, inspired by Genetic Improvement. Assured LLMSE applies a series of semantic filters that discard code that fails to meet these twin guarantees. This overcomes the potential problem of LLM's propensity to hallucinate. It allows us to generate code using LLMs, independently of any human. The human plays the role only of final code reviewer, as they would do with code generated by other human engineers. This paper is an outline of the content of the keynote by Mark Harman at the International Workshop on Interpretability, Robustness, and Benchmarking in Neural Software Engineering, Monday 15th April 2024, Lisbon, Portugal.
title Assured LLM-Based Software Engineering
topic Software Engineering
url https://arxiv.org/abs/2402.04380