Revisit Self-Debugging with Self-Generated Tests for Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xiancai, Tao, Zhengwei, Zhang, Kechi, Zhou, Changzhi, Gu, Wanli, He, Yuanpeng, Zhang, Mengdi, Cai, Xunliang, Zhao, Haiyan, Jin, Zhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909463836884992
author Chen, Xiancai
Tao, Zhengwei
Zhang, Kechi
Zhou, Changzhi
Gu, Wanli
He, Yuanpeng
Zhang, Mengdi
Cai, Xunliang
Zhao, Haiyan
Jin, Zhi
author_facet Chen, Xiancai
Tao, Zhengwei
Zhang, Kechi
Zhou, Changzhi
Gu, Wanli
He, Yuanpeng
Zhang, Mengdi
Cai, Xunliang
Zhao, Haiyan
Jin, Zhi
contents Large language models (LLMs) have shown significant advancements in code generation, but still face challenges on tasks beyond their basic capabilities. Recently, the notion of self-debugging has been proposed to boost the performance of code generation by leveraging execution feedback from tests. Despite its promise, the availability of high-quality tests in real-world scenarios is limited. In this context, self-debugging with self-generated tests is a promising solution but lacks a full exploration of its limitations and practical potential. Therefore, we investigate its efficacy on diverse programming problems. To deepen our understanding, we propose two distinct paradigms for the process: post-execution and in-execution self-debugging. Within the scope of self-contained Python programming tasks, we find that post-execution self-debugging struggles on basic problems but shows potential for improvement on competitive ones, due to the bias introduced by self-generated tests. On the other hand, in-execution self-debugging enables LLMs to mitigate the bias by solely leveraging intermediate states during execution, thereby enhancing code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisit Self-Debugging with Self-Generated Tests for Code Generation
Chen, Xiancai
Tao, Zhengwei
Zhang, Kechi
Zhou, Changzhi
Gu, Wanli
He, Yuanpeng
Zhang, Mengdi
Cai, Xunliang
Zhao, Haiyan
Jin, Zhi
Software Engineering
Artificial Intelligence
Large language models (LLMs) have shown significant advancements in code generation, but still face challenges on tasks beyond their basic capabilities. Recently, the notion of self-debugging has been proposed to boost the performance of code generation by leveraging execution feedback from tests. Despite its promise, the availability of high-quality tests in real-world scenarios is limited. In this context, self-debugging with self-generated tests is a promising solution but lacks a full exploration of its limitations and practical potential. Therefore, we investigate its efficacy on diverse programming problems. To deepen our understanding, we propose two distinct paradigms for the process: post-execution and in-execution self-debugging. Within the scope of self-contained Python programming tasks, we find that post-execution self-debugging struggles on basic problems but shows potential for improvement on competitive ones, due to the bias introduced by self-generated tests. On the other hand, in-execution self-debugging enables LLMs to mitigate the bias by solely leveraging intermediate states during execution, thereby enhancing code generation.
title Revisit Self-Debugging with Self-Generated Tests for Code Generation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2501.12793