Improving the Readability of Automatically Generated Tests using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Biagiola, Matteo, Ghislotti, Gianluca, Tonella, Paolo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912423504510976
author Biagiola, Matteo
Ghislotti, Gianluca
Tonella, Paolo
author_facet Biagiola, Matteo
Ghislotti, Gianluca
Tonella, Paolo
contents Search-based test generators are effective at producing unit tests with high coverage. However, such automatically generated tests have no meaningful test and variable names, making them hard to understand and interpret by developers. On the other hand, large language models (LLMs) can generate highly readable test cases, but they are not able to match the effectiveness of search-based generators, in terms of achieved code coverage. In this paper, we propose to combine the effectiveness of search-based generators with the readability of LLM generated tests. Our approach focuses on improving test and variable names produced by search-based tools, while keeping their semantics (i.e., their coverage) unchanged. Our evaluation on nine industrial and open source LLMs show that our readability improvement transformations are overall semantically-preserving and stable across multiple repetitions. Moreover, a human study with ten professional developers, show that our LLM-improved tests are as readable as developer-written tests, regardless of the LLM employed.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18843
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving the Readability of Automatically Generated Tests using Large Language Models
Biagiola, Matteo
Ghislotti, Gianluca
Tonella, Paolo
Software Engineering
Search-based test generators are effective at producing unit tests with high coverage. However, such automatically generated tests have no meaningful test and variable names, making them hard to understand and interpret by developers. On the other hand, large language models (LLMs) can generate highly readable test cases, but they are not able to match the effectiveness of search-based generators, in terms of achieved code coverage. In this paper, we propose to combine the effectiveness of search-based generators with the readability of LLM generated tests. Our approach focuses on improving test and variable names produced by search-based tools, while keeping their semantics (i.e., their coverage) unchanged. Our evaluation on nine industrial and open source LLMs show that our readability improvement transformations are overall semantically-preserving and stable across multiple repetitions. Moreover, a human study with ten professional developers, show that our LLM-improved tests are as readable as developer-written tests, regardless of the LLM employed.
title Improving the Readability of Automatically Generated Tests using Large Language Models
topic Software Engineering
url https://arxiv.org/abs/2412.18843