Teaching Language Models to Critique via Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Zhihui, Chen, Jie, Chen, Liyu, Mao, Weichao, Xu, Jingjing, Kong, Lingpeng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908680347189248
author Xie, Zhihui
Chen, Jie
Chen, Liyu
Mao, Weichao
Xu, Jingjing
Kong, Lingpeng
author_facet Xie, Zhihui
Chen, Jie
Chen, Liyu
Mao, Weichao
Xu, Jingjing
Kong, Lingpeng
contents Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide accurate judgments and actionable suggestions. In this work, we study LLM critics for code generation and propose $\texttt{CTRL}$, a framework for $\texttt{C}$ritic $\texttt{T}$raining via $\texttt{R}$einforcement $\texttt{L}$earning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with $\texttt{CTRL}$ significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models. Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1% relative improvements across challenging code generation benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Teaching Language Models to Critique via Reinforcement Learning
Xie, Zhihui
Chen, Jie
Chen, Liyu
Mao, Weichao
Xu, Jingjing
Kong, Lingpeng
Machine Learning
Artificial Intelligence
Computation and Language
Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide accurate judgments and actionable suggestions. In this work, we study LLM critics for code generation and propose $\texttt{CTRL}$, a framework for $\texttt{C}$ritic $\texttt{T}$raining via $\texttt{R}$einforcement $\texttt{L}$earning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with $\texttt{CTRL}$ significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models. Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1% relative improvements across challenging code generation benchmarks.
title Teaching Language Models to Critique via Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.03492