CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kapadnis, Manav Nitin, Naik, Atharva, Rose, Carolyn
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916770421407744
author Kapadnis, Manav Nitin
Naik, Atharva
Rose, Carolyn
author_facet Kapadnis, Manav Nitin
Naik, Atharva
Rose, Carolyn
contents Reinforcement learning (RL) to improve code review comment generation requires handling unstructured outputs, making reinforcement learning (RL) feedback challenging. The two main RL approaches, namely RL with Verifiable Feedback (RLVR) and RL with AI Feedback (RLAIF), offer trade-offs: RLVR provides reliable feedback for structured tasks like code generation, while RLAIF works for unstructured outputs but is subjective. We bridge this gap with CRScore++, an RL framework that leverages both LLM-based subjective feedback and verifiable signals for training. Extending CRScore, a code review evaluation metric integrating LLMs with verifiers like linters and code smell detectors, CRScore++ transforms these signals into training rewards. We show that CRScore++ improves a weaker student model through a combination of supervised fine-tuning and RL critique from a stronger teacher model, thus enabling generalization to novel programming languages.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
Kapadnis, Manav Nitin
Naik, Atharva
Rose, Carolyn
Software Engineering
Reinforcement learning (RL) to improve code review comment generation requires handling unstructured outputs, making reinforcement learning (RL) feedback challenging. The two main RL approaches, namely RL with Verifiable Feedback (RLVR) and RL with AI Feedback (RLAIF), offer trade-offs: RLVR provides reliable feedback for structured tasks like code generation, while RLAIF works for unstructured outputs but is subjective. We bridge this gap with CRScore++, an RL framework that leverages both LLM-based subjective feedback and verifiable signals for training. Extending CRScore, a code review evaluation metric integrating LLMs with verifiers like linters and code smell detectors, CRScore++ transforms these signals into training rewards. We show that CRScore++ improves a weaker student model through a combination of supervised fine-tuning and RL critique from a stronger teacher model, thus enabling generalization to novel programming languages.
title CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
topic Software Engineering
url https://arxiv.org/abs/2506.00296