TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wenhan, Yang, Chenyuan, Wang, Zhijie, Huang, Yuheng, Chu, Zhaoyang, Song, Da, Zhang, Lingming, Chen, An Ran, Ma, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917907971178496
author Wang, Wenhan
Yang, Chenyuan
Wang, Zhijie
Huang, Yuheng
Chu, Zhaoyang
Song, Da
Zhang, Lingming
Chen, An Ran
Ma, Lei
author_facet Wang, Wenhan
Yang, Chenyuan
Wang, Zhijie
Huang, Yuheng
Chu, Zhaoyang
Song, Da
Zhang, Lingming
Chen, An Ran
Ma, Lei
contents Testing plays a crucial role in the software development cycle, enabling the detection of bugs, vulnerabilities, and other undesirable behaviors. To perform software testing, testers need to write code snippets that execute the program under test. Recently, researchers have recognized the potential of large language models (LLMs) in software testing. However, there remains a lack of fair comparisons between different LLMs in terms of test case generation capabilities. In this paper, we propose TESTEVAL, a novel benchmark for test case generation with LLMs. We collect 210 Python programs from an online programming platform, LeetCode, and design three different tasks: overall coverage, targeted line/branch coverage, and targeted path coverage. We further evaluate sixteen popular LLMs, including both commercial and open-source ones, on TESTEVAL. We find that generating test cases to cover specific program lines/branches/paths is still challenging for current LLMs, indicating a lack of ability to comprehend program logic and execution paths. We have open-sourced our dataset and benchmark pipelines at https://github.com/LLM4SoftwareTesting/TestEval.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04531
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TESTEVAL: Benchmarking Large Language Models for Test Case Generation
Wang, Wenhan
Yang, Chenyuan
Wang, Zhijie
Huang, Yuheng
Chu, Zhaoyang
Song, Da
Zhang, Lingming
Chen, An Ran
Ma, Lei
Software Engineering
Testing plays a crucial role in the software development cycle, enabling the detection of bugs, vulnerabilities, and other undesirable behaviors. To perform software testing, testers need to write code snippets that execute the program under test. Recently, researchers have recognized the potential of large language models (LLMs) in software testing. However, there remains a lack of fair comparisons between different LLMs in terms of test case generation capabilities. In this paper, we propose TESTEVAL, a novel benchmark for test case generation with LLMs. We collect 210 Python programs from an online programming platform, LeetCode, and design three different tasks: overall coverage, targeted line/branch coverage, and targeted path coverage. We further evaluate sixteen popular LLMs, including both commercial and open-source ones, on TESTEVAL. We find that generating test cases to cover specific program lines/branches/paths is still challenging for current LLMs, indicating a lack of ability to comprehend program logic and execution paths. We have open-sourced our dataset and benchmark pipelines at https://github.com/LLM4SoftwareTesting/TestEval.
title TESTEVAL: Benchmarking Large Language Models for Test Case Generation
topic Software Engineering
url https://arxiv.org/abs/2406.04531