RefineCoder: Iterative Improving of Large Language Models via Adaptive Critique Refinement for Code Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Changzhi, Zhang, Xinyu, Song, Dandan, Chen, Xiancai, Gu, Wanli, Ma, Huipeng, Tian, Yuhang, Zhang, Mengdi, Hu, Linmei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915455342477312
author Zhou, Changzhi
Zhang, Xinyu
Song, Dandan
Chen, Xiancai
Gu, Wanli
Ma, Huipeng
Tian, Yuhang
Zhang, Mengdi
Hu, Linmei
author_facet Zhou, Changzhi
Zhang, Xinyu
Song, Dandan
Chen, Xiancai
Gu, Wanli
Ma, Huipeng
Tian, Yuhang
Zhang, Mengdi
Hu, Linmei
contents Code generation has attracted increasing attention with the rise of Large Language Models (LLMs). Many studies have developed powerful code LLMs by synthesizing code-related instruction data and applying supervised fine-tuning. However, these methods are limited by teacher model distillation and ignore the potential of iterative refinement by self-generated code. In this paper, we propose Adaptive Critique Refinement (ACR), which enables the model to refine itself by self-generated code and external critique, rather than directly imitating the code responses of the teacher model. Concretely, ACR includes a composite scoring system with LLM-as-a-Judge to evaluate the quality of code responses and a selective critique strategy with LLM-as-a-Critic to critique self-generated low-quality code responses. We develop the RefineCoder series by iteratively applying ACR, achieving continuous performance improvement on multiple code generation benchmarks. Compared to the baselines of the same size, our proposed RefineCoder series can achieve comparable or even superior performance using less data.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefineCoder: Iterative Improving of Large Language Models via Adaptive Critique Refinement for Code Generation
Zhou, Changzhi
Zhang, Xinyu
Song, Dandan
Chen, Xiancai
Gu, Wanli
Ma, Huipeng
Tian, Yuhang
Zhang, Mengdi
Hu, Linmei
Computation and Language
Artificial Intelligence
Code generation has attracted increasing attention with the rise of Large Language Models (LLMs). Many studies have developed powerful code LLMs by synthesizing code-related instruction data and applying supervised fine-tuning. However, these methods are limited by teacher model distillation and ignore the potential of iterative refinement by self-generated code. In this paper, we propose Adaptive Critique Refinement (ACR), which enables the model to refine itself by self-generated code and external critique, rather than directly imitating the code responses of the teacher model. Concretely, ACR includes a composite scoring system with LLM-as-a-Judge to evaluate the quality of code responses and a selective critique strategy with LLM-as-a-Critic to critique self-generated low-quality code responses. We develop the RefineCoder series by iteratively applying ACR, achieving continuous performance improvement on multiple code generation benchmarks. Compared to the baselines of the same size, our proposed RefineCoder series can achieve comparable or even superior performance using less data.
title RefineCoder: Iterative Improving of Large Language Models via Adaptive Critique Refinement for Code Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.09183