Evaluating Large Language Models for the Generation of Unit Tests with Equivalence Partitions and Boundary Values

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rodríguez, Martín, Rossi, Gustavo, Fernandez, Alejandro
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913838490714112
author Rodríguez, Martín
Rossi, Gustavo
Fernandez, Alejandro
author_facet Rodríguez, Martín
Rossi, Gustavo
Fernandez, Alejandro
contents The design and implementation of unit tests is a complex task many programmers neglect. This research evaluates the potential of Large Language Models (LLMs) in automatically generating test cases, comparing them with manual tests. An optimized prompt was developed, that integrates code and requirements, covering critical cases such as equivalence partitions and boundary values. The strengths and weaknesses of LLMs versus trained programmers were compared through quantitative metrics and manual qualitative analysis. The results show that the effectiveness of LLMs depends on well-designed prompts, robust implementation, and precise requirements. Although flexible and promising, LLMs still require human supervision. This work highlights the importance of manual qualitative analysis as an essential complement to automation in unit test evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09830
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Large Language Models for the Generation of Unit Tests with Equivalence Partitions and Boundary Values
Rodríguez, Martín
Rossi, Gustavo
Fernandez, Alejandro
Software Engineering
Artificial Intelligence
The design and implementation of unit tests is a complex task many programmers neglect. This research evaluates the potential of Large Language Models (LLMs) in automatically generating test cases, comparing them with manual tests. An optimized prompt was developed, that integrates code and requirements, covering critical cases such as equivalence partitions and boundary values. The strengths and weaknesses of LLMs versus trained programmers were compared through quantitative metrics and manual qualitative analysis. The results show that the effectiveness of LLMs depends on well-designed prompts, robust implementation, and precise requirements. Although flexible and promising, LLMs still require human supervision. This work highlights the importance of manual qualitative analysis as an essential complement to automation in unit test evaluation.
title Evaluating Large Language Models for the Generation of Unit Tests with Equivalence Partitions and Boundary Values
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2505.09830