GenX: Mastering Code and Test Generation with Execution Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Nan, Liu, Yafei, Chen, Chen, Lu, Haonan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912160285720576
author Wang, Nan
Liu, Yafei
Chen, Chen
Lu, Haonan
author_facet Wang, Nan
Liu, Yafei
Chen, Chen
Lu, Haonan
contents Recent advancements in language modeling have enabled the translation of natural language into code, and the use of execution feedback to improve code generation. However, these methods often rely heavily on pre-existing test cases, which may not always be available or comprehensive. In this work, we propose a novel approach that concurrently trains a code generation model and a test generation model, utilizing execution feedback to refine and enhance the performance of both. We introduce two strategies for test and code data augmentation and a new scoring function for code and test ranking. We experiment on the APPS dataset and demonstrate that our approach can effectively generate and augment test cases, filter and synthesize correct code solutions, and rank the quality of generated code and tests. The results demonstrate that our models, when iteratively trained with an increasing number of test cases and code solutions, outperform those trained on the original dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13464
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenX: Mastering Code and Test Generation with Execution Feedback
Wang, Nan
Liu, Yafei
Chen, Chen
Lu, Haonan
Software Engineering
Computation and Language
Recent advancements in language modeling have enabled the translation of natural language into code, and the use of execution feedback to improve code generation. However, these methods often rely heavily on pre-existing test cases, which may not always be available or comprehensive. In this work, we propose a novel approach that concurrently trains a code generation model and a test generation model, utilizing execution feedback to refine and enhance the performance of both. We introduce two strategies for test and code data augmentation and a new scoring function for code and test ranking. We experiment on the APPS dataset and demonstrate that our approach can effectively generate and augment test cases, filter and synthesize correct code solutions, and rank the quality of generated code and tests. The results demonstrate that our models, when iteratively trained with an increasing number of test cases and code solutions, outperform those trained on the original dataset.
title GenX: Mastering Code and Test Generation with Execution Feedback
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2412.13464