Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Zi, Shen, Sheng, Kulikov, Ilia, Shang, Jingbo, Weston, Jason, Nie, Yixin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914363569340416
author Lin, Zi
Shen, Sheng
Kulikov, Ilia
Shang, Jingbo
Weston, Jason
Nie, Yixin
author_facet Lin, Zi
Shen, Sheng
Kulikov, Ilia
Shang, Jingbo
Weston, Jason
Nie, Yixin
contents Recent advances in large language models (LLMs) have improved their performance on coding benchmarks. However, improvement is plateauing due to the exhaustion of readily available high-quality data. Prior work has shown the potential of synthetic self-instruct data, but naively training on a model's own outputs can cause error accumulation, especially in coding tasks, where generalization may collapse due to overly simple or erroneous training data, highlighting the need for rigorous quality checks on synthetic data. In this work, we explore an effective approach whereby the model itself verifies the correctness of its own data. We thus propose Sol-Ver, a self-play solver-verifier framework that jointly improves a single model's code and test generation capacity. By iteratively refining code (LLM-as-a-solver) and tests (LLM-as-a-verifier) together, we boost both capabilities without relying on human annotations or larger teacher models. Experiments with the Llama 3.1 8B model demonstrate substantial performance enhancements, achieving average relative improvements of 19.63% in code generation and 17.49% in test generation on MBPP and LiveCodeBench.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
Lin, Zi
Shen, Sheng
Kulikov, Ilia
Shang, Jingbo
Weston, Jason
Nie, Yixin
Software Engineering
Recent advances in large language models (LLMs) have improved their performance on coding benchmarks. However, improvement is plateauing due to the exhaustion of readily available high-quality data. Prior work has shown the potential of synthetic self-instruct data, but naively training on a model's own outputs can cause error accumulation, especially in coding tasks, where generalization may collapse due to overly simple or erroneous training data, highlighting the need for rigorous quality checks on synthetic data. In this work, we explore an effective approach whereby the model itself verifies the correctness of its own data. We thus propose Sol-Ver, a self-play solver-verifier framework that jointly improves a single model's code and test generation capacity. By iteratively refining code (LLM-as-a-solver) and tests (LLM-as-a-verifier) together, we boost both capabilities without relying on human annotations or larger teacher models. Experiments with the Llama 3.1 8B model demonstrate substantial performance enhancements, achieving average relative improvements of 19.63% in code generation and 17.49% in test generation on MBPP and LiveCodeBench.
title Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
topic Software Engineering
url https://arxiv.org/abs/2502.14948