Large Language Model Guided Self-Debugging Code Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Adnan, Muntasir, Xu, Zhiwei, Kuhn, Carlos C. N.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911017874751488
author Adnan, Muntasir
Xu, Zhiwei
Kuhn, Carlos C. N.
author_facet Adnan, Muntasir
Xu, Zhiwei
Kuhn, Carlos C. N.
contents Automated code generation is gaining significant importance in intelligent computer programming and system deployment. However, current approaches often face challenges in computational efficiency and lack robust mechanisms for code parsing and error correction. In this work, we propose a novel framework, PyCapsule, with a simple yet effective two-agent pipeline and efficient self-debugging modules for Python code generation. PyCapsule features sophisticated prompt inference, iterative error handling, and case testing, ensuring high generation stability, safety, and correctness. Empirically, PyCapsule achieves up to 5.7% improvement of success rate on HumanEval, 10.3% on HumanEval-ET, and 24.4% on BigCodeBench compared to the state-of-art methods. We also observe a decrease in normalized success rate given more self-debugging attempts, potentially affected by limited and noisy error feedback in retention. PyCapsule demonstrates broader impacts on advancing lightweight and efficient code generation for artificial intelligence systems.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Model Guided Self-Debugging Code Generation
Adnan, Muntasir
Xu, Zhiwei
Kuhn, Carlos C. N.
Software Engineering
Artificial Intelligence
Automated code generation is gaining significant importance in intelligent computer programming and system deployment. However, current approaches often face challenges in computational efficiency and lack robust mechanisms for code parsing and error correction. In this work, we propose a novel framework, PyCapsule, with a simple yet effective two-agent pipeline and efficient self-debugging modules for Python code generation. PyCapsule features sophisticated prompt inference, iterative error handling, and case testing, ensuring high generation stability, safety, and correctness. Empirically, PyCapsule achieves up to 5.7% improvement of success rate on HumanEval, 10.3% on HumanEval-ET, and 24.4% on BigCodeBench compared to the state-of-art methods. We also observe a decrease in normalized success rate given more self-debugging attempts, potentially affected by limited and noisy error feedback in retention. PyCapsule demonstrates broader impacts on advancing lightweight and efficient code generation for artificial intelligence systems.
title Large Language Model Guided Self-Debugging Code Generation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2502.02928