CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jiacheng, Pang, Bo, Qu, Jin, Hayashi, Hiroaki, Xiong, Caiming, Zhou, Yingbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909490559844352
author Xu, Jiacheng
Pang, Bo
Qu, Jin
Hayashi, Hiroaki
Xiong, Caiming
Zhou, Yingbo
author_facet Xu, Jiacheng
Pang, Bo
Qu, Jin
Hayashi, Hiroaki
Xiong, Caiming
Zhou, Yingbo
contents Software testing is a critical aspect of software development, yet generating test cases remains a routine task for engineers. This paper presents a benchmark, CLOVER, to evaluate models' capabilities in generating and completing test cases under specific conditions. Spanning from simple assertion completions to writing test cases that cover specific code blocks across multiple files, these tasks are based on 12 python repositories, analyzing 845 problems with context lengths ranging from 4k to 128k tokens. Utilizing code testing frameworks, we propose a method to construct retrieval contexts using coverage information. While models exhibit comparable performance with short contexts, notable differences emerge with 16k contexts. Notably, models like GPT-4o and Claude 3.5 can effectively leverage relevant snippets; however, all models score below 35\% on the complex Task III, even with the oracle context provided, underscoring the benchmark's significance and the potential for model improvement. The benchmark is containerized for code execution across tasks, and we will release the code, data, and construction methodologies.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
Xu, Jiacheng
Pang, Bo
Qu, Jin
Hayashi, Hiroaki
Xiong, Caiming
Zhou, Yingbo
Software Engineering
Artificial Intelligence
Machine Learning
Software testing is a critical aspect of software development, yet generating test cases remains a routine task for engineers. This paper presents a benchmark, CLOVER, to evaluate models' capabilities in generating and completing test cases under specific conditions. Spanning from simple assertion completions to writing test cases that cover specific code blocks across multiple files, these tasks are based on 12 python repositories, analyzing 845 problems with context lengths ranging from 4k to 128k tokens. Utilizing code testing frameworks, we propose a method to construct retrieval contexts using coverage information. While models exhibit comparable performance with short contexts, notable differences emerge with 16k contexts. Notably, models like GPT-4o and Claude 3.5 can effectively leverage relevant snippets; however, all models score below 35\% on the complex Task III, even with the oracle context provided, underscoring the benchmark's significance and the potential for model improvement. The benchmark is containerized for code execution across tasks, and we will release the code, data, and construction methodologies.
title CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
topic Software Engineering
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.08806