Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Yuhan, Chen, Zian, Quan, Kun, Zhang, Ziliang, Wang, Yu, Dong, Xiaoning, Feng, Yeqi, He, Guanzhong, Huang, Jingcheng, Li, Jianhao, Tan, Yixuan, Tang, Jiafu, Tang, Yilin, Wu, Junlei, Xiao, Qianyu, Zheng, Can, Zhou, Shouchen, Zhu, Yuxiang, Huang, Yiming, He, Tianxing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911375664611328
author Cao, Yuhan
Chen, Zian
Quan, Kun
Zhang, Ziliang
Wang, Yu
Dong, Xiaoning
Feng, Yeqi
He, Guanzhong
Huang, Jingcheng
Li, Jianhao
Tan, Yixuan
Tang, Jiafu
Tang, Yilin
Wu, Junlei
Xiao, Qianyu
Zheng, Can
Zhou, Shouchen
Zhu, Yuxiang
Huang, Yiming
He, Tianxing
author_facet Cao, Yuhan
Chen, Zian
Quan, Kun
Zhang, Ziliang
Wang, Yu
Dong, Xiaoning
Feng, Yeqi
He, Guanzhong
Huang, Jingcheng
Li, Jianhao
Tan, Yixuan
Tang, Jiafu
Tang, Yilin
Wu, Junlei
Xiao, Qianyu
Zheng, Can
Zhou, Shouchen
Zhu, Yuxiang
Huang, Yiming
He, Tianxing
contents Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, capable of tackling complex tasks during inference. However, the extent to which LLMs can be utilized for code checking or debugging through test case generation remains largely unexplored. We investigate this problem from the perspective of competition-level programming (CP) programs and propose TCGBench, a Benchmark for (LLM generation of) Test Case Generators. This benchmark comprises two tasks, aimed at studying the capabilities of LLMs in (1) generating valid test case generators for a given CP problem, and further (2) generating targeted test case generators that expose bugs in human-written code. Experimental results indicate that while state-of-the-art LLMs can generate valid test case generators in most cases, most LLMs struggle to generate targeted test cases that reveal flaws in human code effectively. Especially, even advanced reasoning models (e.g., o3-mini) fall significantly short of human performance in the task of generating targeted generators. Furthermore, we construct a high-quality, manually curated dataset of instructions for generating targeted generators. Analysis demonstrates that the performance of LLMs can be enhanced with the aid of this dataset, by both prompting and fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
Cao, Yuhan
Chen, Zian
Quan, Kun
Zhang, Ziliang
Wang, Yu
Dong, Xiaoning
Feng, Yeqi
He, Guanzhong
Huang, Jingcheng
Li, Jianhao
Tan, Yixuan
Tang, Jiafu
Tang, Yilin
Wu, Junlei
Xiao, Qianyu
Zheng, Can
Zhou, Shouchen
Zhu, Yuxiang
Huang, Yiming
He, Tianxing
Computation and Language
Artificial Intelligence
Software Engineering
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, capable of tackling complex tasks during inference. However, the extent to which LLMs can be utilized for code checking or debugging through test case generation remains largely unexplored. We investigate this problem from the perspective of competition-level programming (CP) programs and propose TCGBench, a Benchmark for (LLM generation of) Test Case Generators. This benchmark comprises two tasks, aimed at studying the capabilities of LLMs in (1) generating valid test case generators for a given CP problem, and further (2) generating targeted test case generators that expose bugs in human-written code. Experimental results indicate that while state-of-the-art LLMs can generate valid test case generators in most cases, most LLMs struggle to generate targeted test cases that reveal flaws in human code effectively. Especially, even advanced reasoning models (e.g., o3-mini) fall significantly short of human performance in the task of generating targeted generators. Furthermore, we construct a high-quality, manually curated dataset of instructions for generating targeted generators. Analysis demonstrates that the performance of LLMs can be enhanced with the aid of this dataset, by both prompting and fine-tuning.
title Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2506.06821