Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Jia, Yang, Xinyu, Zhang, Hongzhi, Liu, Yahui, Zhang, Jingyuan, Wang, Qi, Zhang, Fuzheng, Zhou, Guorui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911148825116672
author Fu, Jia
Yang, Xinyu
Zhang, Hongzhi
Liu, Yahui
Zhang, Jingyuan
Wang, Qi
Zhang, Fuzheng
Zhou, Guorui
author_facet Fu, Jia
Yang, Xinyu
Zhang, Hongzhi
Liu, Yahui
Zhang, Jingyuan
Wang, Qi
Zhang, Fuzheng
Zhou, Guorui
contents Precise, correct feedback is crucial for effectively training large language models (LLMs) in code reinforcement learning. However, synthesizing high-quality test cases remains a profoundly challenging and unsolved problem. In this work, we present Klear-CodeTest, a comprehensive test case synthesis framework featuring rigorous verification to ensure quality and reliability of test cases. Our approach achieves broad coverage of programming problems via a novel Generator-Validation (G-V) framework, ensuring correctness through a consistency validation mechanism that verifies outputs against gold solutions. The proposed G-V framework generates comprehensive test cases including both regular and corner cases, enhancing test coverage and discriminative power for solution correctness assessment in code reinforcement learning. In addition, we design a multi-layered security sandbox system optimized for online verification platforms, guaranteeing safe and reliable code execution. Through comprehensive experiments, we demonstrate the effectiveness of our curated dataset, showing significant improvements in model performance and training stability. The source codes, curated dataset and sandbox system are available at: https://github.com/Kwai-Klear/CodeTest.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05710
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
Fu, Jia
Yang, Xinyu
Zhang, Hongzhi
Liu, Yahui
Zhang, Jingyuan
Wang, Qi
Zhang, Fuzheng
Zhou, Guorui
Software Engineering
Artificial Intelligence
Precise, correct feedback is crucial for effectively training large language models (LLMs) in code reinforcement learning. However, synthesizing high-quality test cases remains a profoundly challenging and unsolved problem. In this work, we present Klear-CodeTest, a comprehensive test case synthesis framework featuring rigorous verification to ensure quality and reliability of test cases. Our approach achieves broad coverage of programming problems via a novel Generator-Validation (G-V) framework, ensuring correctness through a consistency validation mechanism that verifies outputs against gold solutions. The proposed G-V framework generates comprehensive test cases including both regular and corner cases, enhancing test coverage and discriminative power for solution correctness assessment in code reinforcement learning. In addition, we design a multi-layered security sandbox system optimized for online verification platforms, guaranteeing safe and reliable code execution. Through comprehensive experiments, we demonstrate the effectiveness of our curated dataset, showing significant improvements in model performance and training stability. The source codes, curated dataset and sandbox system are available at: https://github.com/Kwai-Klear/CodeTest.
title Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.05710