Reward Modeling with Weak Supervision for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hauptvogel, Ben, Ostendorff, Malte, Rehm, Georg, Möller, Sebastian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913565769728000
author Hauptvogel, Ben
Ostendorff, Malte
Rehm, Georg
Möller, Sebastian
author_facet Hauptvogel, Ben
Ostendorff, Malte
Rehm, Georg
Möller, Sebastian
contents Recent advancements in large language models (LLMs) have led to their increased application across various tasks, with reinforcement learning from human feedback (RLHF) being a crucial part of their training to align responses with user intentions. In the RLHF process, a reward model is trained using responses preferences determined by human labelers or AI systems, which then refines the LLM through reinforcement learning. This work introduces weak supervision as a strategy to extend RLHF datasets and enhance reward model performance. Weak supervision employs noisy or imprecise data labeling, reducing reliance on expensive manually labeled data. By analyzing RLHF datasets to identify heuristics that correlate with response preference, we wrote simple labeling functions and then calibrated a label model to weakly annotate unlabeled data. Our evaluation show that while weak supervision significantly benefits smaller datasets by improving reward model performance, its effectiveness decreases with larger, originally labeled datasets. Additionally, using an LLM to generate and then weakly label responses offers a promising method for extending preference data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20869
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reward Modeling with Weak Supervision for Language Models
Hauptvogel, Ben
Ostendorff, Malte
Rehm, Georg
Möller, Sebastian
Computation and Language
Recent advancements in large language models (LLMs) have led to their increased application across various tasks, with reinforcement learning from human feedback (RLHF) being a crucial part of their training to align responses with user intentions. In the RLHF process, a reward model is trained using responses preferences determined by human labelers or AI systems, which then refines the LLM through reinforcement learning. This work introduces weak supervision as a strategy to extend RLHF datasets and enhance reward model performance. Weak supervision employs noisy or imprecise data labeling, reducing reliance on expensive manually labeled data. By analyzing RLHF datasets to identify heuristics that correlate with response preference, we wrote simple labeling functions and then calibrated a label model to weakly annotate unlabeled data. Our evaluation show that while weak supervision significantly benefits smaller datasets by improving reward model performance, its effectiveness decreases with larger, originally labeled datasets. Additionally, using an LLM to generate and then weakly label responses offers a promising method for extending preference data.
title Reward Modeling with Weak Supervision for Language Models
topic Computation and Language
url https://arxiv.org/abs/2410.20869