Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taherkhani, Hamed, Shin, Jiho, Tahir, Muhammad Ammar, Misu, Md Rakib Hossain, Gattani, Vineet Sunil, Hemmati, Hadi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915816868413440
author Taherkhani, Hamed
Shin, Jiho
Tahir, Muhammad Ammar
Misu, Md Rakib Hossain
Gattani, Vineet Sunil
Hemmati, Hadi
author_facet Taherkhani, Hamed
Shin, Jiho
Tahir, Muhammad Ammar
Misu, Md Rakib Hossain
Gattani, Vineet Sunil
Hemmati, Hadi
contents Modern Large Language Model (LLM)-based programming agents often rely on test execution feedback to refine their generated code. These tests are synthetically generated by LLMs. However, LLMs may produce invalid or hallucinated test cases, which can mislead feedback loops and degrade the performance of agents in refining and improving code. This paper introduces VALTEST, a novel framework that leverages semantic entropy to automatically validate test cases generated by LLMs. Analyzing the semantic structure of test cases and computing entropy-based uncertainty measures, VALTEST trains a machine learning model to classify test cases as valid or invalid and filters out invalid test cases. Experiments on multiple benchmark datasets and various LLMs show that VALTEST not only boosts test validity by up to 29% but also improves code generation performance, as evidenced by significant increases in pass@1 scores. Our extensive experiments also reveal that semantic entropy is a reliable indicator to distinguish between valid and invalid test cases, which provides a robust solution for improving the correctness of LLM-generated test cases used in software testing and code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08254
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
Taherkhani, Hamed
Shin, Jiho
Tahir, Muhammad Ammar
Misu, Md Rakib Hossain
Gattani, Vineet Sunil
Hemmati, Hadi
Software Engineering
Artificial Intelligence
Modern Large Language Model (LLM)-based programming agents often rely on test execution feedback to refine their generated code. These tests are synthetically generated by LLMs. However, LLMs may produce invalid or hallucinated test cases, which can mislead feedback loops and degrade the performance of agents in refining and improving code. This paper introduces VALTEST, a novel framework that leverages semantic entropy to automatically validate test cases generated by LLMs. Analyzing the semantic structure of test cases and computing entropy-based uncertainty measures, VALTEST trains a machine learning model to classify test cases as valid or invalid and filters out invalid test cases. Experiments on multiple benchmark datasets and various LLMs show that VALTEST not only boosts test validity by up to 29% but also improves code generation performance, as evidenced by significant increases in pass@1 scores. Our extensive experiments also reveal that semantic entropy is a reliable indicator to distinguish between valid and invalid test cases, which provides a robust solution for improving the correctness of LLM-generated test cases used in software testing and code generation.
title Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2411.08254