DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Ruiyi, Qin, Peijia, Cao, Qi, Xie, Pengtao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909966691991552
author Zhang, Ruiyi
Qin, Peijia
Cao, Qi
Xie, Pengtao
author_facet Zhang, Ruiyi
Qin, Peijia
Cao, Qi
Xie, Pengtao
contents Process Reward Models (PRMs) have become essential for improving Large Language Models (LLMs) via test-time scaling, yet their effectiveness in coding remains limited due to the lack of meaningful step decompositions in code and the noise of Monte-Carlo-generated partial labels. We propose DreamPRM-Code, a coding-focused PRM that treats functions as reasoning steps using a Chain-of-Function prompting strategy to induce modular code generation, enabling PRM training and application analogous to mathematical reasoning tasks. To address label noise, DreamPRM-Code introduces a meta-learning-based correction mechanism that leverages clean final-solution unit-test labels and performs bi-level optimization to refine intermediate labels. Applying on test-time scaling, DreamPRM-Code achieved state-of-the-art performance on LiveCodeBench with 80.9 pass@1 rate, surpassing OpenAI o4-mini.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15000
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
Zhang, Ruiyi
Qin, Peijia
Cao, Qi
Xie, Pengtao
Machine Learning
Artificial Intelligence
Computation and Language
Process Reward Models (PRMs) have become essential for improving Large Language Models (LLMs) via test-time scaling, yet their effectiveness in coding remains limited due to the lack of meaningful step decompositions in code and the noise of Monte-Carlo-generated partial labels. We propose DreamPRM-Code, a coding-focused PRM that treats functions as reasoning steps using a Chain-of-Function prompting strategy to induce modular code generation, enabling PRM training and application analogous to mathematical reasoning tasks. To address label noise, DreamPRM-Code introduces a meta-learning-based correction mechanism that leverages clean final-solution unit-test labels and performs bi-level optimization to refine intermediate labels. Applying on test-time scaling, DreamPRM-Code achieved state-of-the-art performance on LiveCodeBench with 80.9 pass@1 rate, surpassing OpenAI o4-mini.
title DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.15000