CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gupta, Taneesh, Shandilya, Shivam, Zhang, Xuchao, Madhavan, Rahul, Ghosh, Supriyo, Bansal, Chetan, Yao, Huaxiu, Rajmohan, Saravan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908230360236032
author Gupta, Taneesh
Shandilya, Shivam
Zhang, Xuchao
Madhavan, Rahul
Ghosh, Supriyo
Bansal, Chetan
Yao, Huaxiu
Rajmohan, Saravan
author_facet Gupta, Taneesh
Shandilya, Shivam
Zhang, Xuchao
Madhavan, Rahul
Ghosh, Supriyo
Bansal, Chetan
Yao, Huaxiu
Rajmohan, Saravan
contents Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In reinforcement learning from human feedback (RLHF) and more generally during post-training flawed reward signals often lead to outputs that optimize for these spurious correlates instead of genuine quality or correctness. We propose Context-Aware Reward Modeling (CARMO), a novel approach that first generates dynamic, context-relevant criteria to ground the reward model before producing reward scores. Unlike prior methods that rely on static rubrics, CARMO leverages large language models (LLMs) to adaptively create evaluation criteria such as logical consistency, clarity, and depth tailored to the user query. Our theoretical analysis shows that such criteria generation can mitigate reward hacking. We further demonstrate that CARMO can be distilled into smaller models, reducing the computational cost of alignment. We establish a new state-of-the-art performance in zero-shot settings for generative models, achieving a 2.1\% improvement on Reward Bench. Furthermore, alignment performed on the CARMO-curated preference dataset achieves 22.5\% and 21.1\% LC-WR and WR, respectively, on Mistral-Base (7B).
format Preprint
id arxiv_https___arxiv_org_abs_2410_21545
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
Gupta, Taneesh
Shandilya, Shivam
Zhang, Xuchao
Madhavan, Rahul
Ghosh, Supriyo
Bansal, Chetan
Yao, Huaxiu
Rajmohan, Saravan
Computation and Language
Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In reinforcement learning from human feedback (RLHF) and more generally during post-training flawed reward signals often lead to outputs that optimize for these spurious correlates instead of genuine quality or correctness. We propose Context-Aware Reward Modeling (CARMO), a novel approach that first generates dynamic, context-relevant criteria to ground the reward model before producing reward scores. Unlike prior methods that rely on static rubrics, CARMO leverages large language models (LLMs) to adaptively create evaluation criteria such as logical consistency, clarity, and depth tailored to the user query. Our theoretical analysis shows that such criteria generation can mitigate reward hacking. We further demonstrate that CARMO can be distilled into smaller models, reducing the computational cost of alignment. We establish a new state-of-the-art performance in zero-shot settings for generative models, achieving a 2.1\% improvement on Reward Bench. Furthermore, alignment performed on the CARMO-curated preference dataset achieves 22.5\% and 21.1\% LC-WR and WR, respectively, on Mistral-Base (7B).
title CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
topic Computation and Language
url https://arxiv.org/abs/2410.21545