Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Mingzhe, Tuan, Luu Anh, Liu, Yue, Qing, Yuhao, Huang, Dong, He, Xinyi, Liu, Qian, Ma, Zejun, Ng, See-kiong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918044059566080
author Du, Mingzhe
Tuan, Luu Anh
Liu, Yue
Qing, Yuhao
Huang, Dong
He, Xinyi
Liu, Qian
Ma, Zejun
Ng, See-kiong
author_facet Du, Mingzhe
Tuan, Luu Anh
Liu, Yue
Qing, Yuhao
Huang, Dong
He, Xinyi
Liu, Qian
Ma, Zejun
Ng, See-kiong
contents Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO). Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both pass@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
Du, Mingzhe
Tuan, Luu Anh
Liu, Yue
Qing, Yuhao
Huang, Dong
He, Xinyi
Liu, Qian
Ma, Zejun
Ng, See-kiong
Software Engineering
Artificial Intelligence
Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO). Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both pass@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency.
title Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2505.23387