RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Reber, David, Richardson, Sean, Nief, Todd, Garbacea, Cristina, Veitch, Victor
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909616510599168
author Reber, David
Richardson, Sean
Nief, Todd
Garbacea, Cristina
Veitch, Victor
author_facet Reber, David
Richardson, Sean
Nief, Todd
Garbacea, Cristina
Veitch, Victor
contents Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective method for measuring the sensitivity of a reward model to high-level attributes of responses, such as sentiment, helpfulness, or complexity. Importantly, RATE measures the causal effect of an attribute on the reward. RATE uses LLMs to rewrite responses to produce imperfect counterfactuals examples that can be used to measure causal effects. A key challenge is that these rewrites are imperfect in a manner that can induce substantial bias in the estimated sensitivity of the reward model to the attribute. The core idea of RATE is to adjust for this imperfect-rewrite effect by rewriting twice. We establish the validity of the RATE procedure and show empirically that it is an effective estimator.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11348
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
Reber, David
Richardson, Sean
Nief, Todd
Garbacea, Cristina
Veitch, Victor
Computation and Language
Artificial Intelligence
Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective method for measuring the sensitivity of a reward model to high-level attributes of responses, such as sentiment, helpfulness, or complexity. Importantly, RATE measures the causal effect of an attribute on the reward. RATE uses LLMs to rewrite responses to produce imperfect counterfactuals examples that can be used to measure causal effects. A key challenge is that these rewrites are imperfect in a manner that can induce substantial bias in the estimated sensitivity of the reward model to the attribute. The core idea of RATE is to adjust for this imperfect-rewrite effect by rewriting twice. We establish the validity of the RATE procedure and show empirically that it is an effective estimator.
title RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.11348