Process Supervision-Guided Policy Optimization for Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Ning, Wu, Zheng, Zheng, Renjie, Wei, Ziyun, Shi, Wenlei, Jin, Xing, Liu, Guanlin, Dun, Chen, Huang, Liang, Yan, Lin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910813091004416
author Dai, Ning
Wu, Zheng
Zheng, Renjie
Wei, Ziyun
Shi, Wenlei
Jin, Xing
Liu, Guanlin
Dun, Chen
Huang, Liang
Yan, Lin
author_facet Dai, Ning
Wu, Zheng
Zheng, Renjie
Wei, Ziyun
Shi, Wenlei
Jin, Xing
Liu, Guanlin
Dun, Chen
Huang, Liang
Yan, Lin
contents Reinforcement learning (RL) with unit test feedback has enhanced large language models' (LLMs) code generation, but relies on sparse rewards provided only after complete code evaluation, limiting learning efficiency and incremental improvements. When generated code fails all unit tests, no learning signal is received, hindering progress on complex tasks. To address this, we propose a Process Reward Model (PRM) that delivers dense, line-level feedback on code correctness during generation, mimicking human code refinement and providing immediate guidance. We explore various strategies for training PRMs and integrating them into the RL framework, finding that using PRMs both as dense rewards and for value function initialization significantly boosts performance. Our experimental results also highlight the effectiveness of PRMs in enhancing RL-driven code generation, especially for long-horizon scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Process Supervision-Guided Policy Optimization for Code Generation
Dai, Ning
Wu, Zheng
Zheng, Renjie
Wei, Ziyun
Shi, Wenlei
Jin, Xing
Liu, Guanlin
Dun, Chen
Huang, Liang
Yan, Lin
Artificial Intelligence
I.2.7,
Reinforcement learning (RL) with unit test feedback has enhanced large language models' (LLMs) code generation, but relies on sparse rewards provided only after complete code evaluation, limiting learning efficiency and incremental improvements. When generated code fails all unit tests, no learning signal is received, hindering progress on complex tasks. To address this, we propose a Process Reward Model (PRM) that delivers dense, line-level feedback on code correctness during generation, mimicking human code refinement and providing immediate guidance. We explore various strategies for training PRMs and integrating them into the RL framework, finding that using PRMs both as dense rewards and for value function initialization significantly boosts performance. Our experimental results also highlight the effectiveness of PRMs in enhancing RL-driven code generation, especially for long-horizon scenarios.
title Process Supervision-Guided Policy Optimization for Code Generation
topic Artificial Intelligence
I.2.7,
url https://arxiv.org/abs/2410.17621