Validating Formal Specifications with LLM-generated Test Cases

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cunha, Alcino, Macedo, Nuno
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918343869464576
author Cunha, Alcino
Macedo, Nuno
author_facet Cunha, Alcino
Macedo, Nuno
contents Validation is a central activity when developing formal specifications. Similarly to coding, a possible validation technique is to define upfront test cases or scenarios that a future specification should satisfy or not. Unfortunately, specifying such test cases is burdensome and error prone, which could cause users to skip this validation task. This paper reports the results of an empirical evaluation of using pre-trained large language models (LLMs) to automate the generation of test cases from natural language requirements. In particular, we focus on test cases for structural requirements of simple domain models formalized in the Alloy specification language. Our evaluation focuses on the state-of-the-art GPT-5 model, but results from other closed- and open-source LLMs are also reported. The results show that, in this context, GPT-5 is already quite effective at generating positive (and negative) test cases that are syntactically correct and that satisfy (or not) the given requirement, and that can detect many wrong specifications written by humans.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Validating Formal Specifications with LLM-generated Test Cases
Cunha, Alcino
Macedo, Nuno
Software Engineering
D.2.1; D.2.4; D.2.5
Validation is a central activity when developing formal specifications. Similarly to coding, a possible validation technique is to define upfront test cases or scenarios that a future specification should satisfy or not. Unfortunately, specifying such test cases is burdensome and error prone, which could cause users to skip this validation task. This paper reports the results of an empirical evaluation of using pre-trained large language models (LLMs) to automate the generation of test cases from natural language requirements. In particular, we focus on test cases for structural requirements of simple domain models formalized in the Alloy specification language. Our evaluation focuses on the state-of-the-art GPT-5 model, but results from other closed- and open-source LLMs are also reported. The results show that, in this context, GPT-5 is already quite effective at generating positive (and negative) test cases that are syntactically correct and that satisfy (or not) the given requirement, and that can detect many wrong specifications written by humans.
title Validating Formal Specifications with LLM-generated Test Cases
topic Software Engineering
D.2.1; D.2.4; D.2.5
url https://arxiv.org/abs/2510.23350