A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: LeVine, Will, Pikus, Benjamin, Chen, Anthony, Hendryx, Sean
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914651066859520
author LeVine, Will
Pikus, Benjamin
Chen, Anthony
Hendryx, Sean
author_facet LeVine, Will
Pikus, Benjamin
Chen, Anthony
Hendryx, Sean
contents Foundation models, specifically Large Language Models (LLMs), have lately gained wide-spread attention and adoption. Reinforcement Learning with Human Feedback (RLHF) involves training a reward model to capture desired behaviors, which is then used to align LLM's. These reward models are additionally used at inference-time to estimate LLM responses' adherence to those desired behaviors. However, there is little work measuring how robust these reward models are to distribution shifts. In this work, we evaluate how reward model performance - measured via accuracy and calibration (i.e. alignment between accuracy and confidence) - is affected by distribution shift. We show novel calibration patterns and accuracy drops due to OOD prompts and responses, and that the reward model is more sensitive to shifts in responses than prompts. Additionally, we adapt an OOD detection technique commonly used in classification to the reward model setting to detect these distribution shifts in prompts and responses.
format Preprint
id arxiv_https___arxiv_org_abs_2311_14743
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
LeVine, Will
Pikus, Benjamin
Chen, Anthony
Hendryx, Sean
Computation and Language
Machine Learning
Foundation models, specifically Large Language Models (LLMs), have lately gained wide-spread attention and adoption. Reinforcement Learning with Human Feedback (RLHF) involves training a reward model to capture desired behaviors, which is then used to align LLM's. These reward models are additionally used at inference-time to estimate LLM responses' adherence to those desired behaviors. However, there is little work measuring how robust these reward models are to distribution shifts. In this work, we evaluate how reward model performance - measured via accuracy and calibration (i.e. alignment between accuracy and confidence) - is affected by distribution shift. We show novel calibration patterns and accuracy drops due to OOD prompts and responses, and that the reward model is more sensitive to shifts in responses than prompts. Additionally, we adapt an OOD detection technique commonly used in classification to the reward model setting to detect these distribution shifts in prompts and responses.
title A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2311.14743