Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hadji-Kyriacou, Avelina Asada, Arandjelovic, Ognjen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914816605552640
author Hadji-Kyriacou, Avelina Asada
Arandjelovic, Ognjen
author_facet Hadji-Kyriacou, Avelina Asada
Arandjelovic, Ognjen
contents Pre-trained Language Models (LMs) exhibit strong zero-shot and in-context learning capabilities; however, their behaviors are often difficult to control. By utilizing Reinforcement Learning from Human Feedback (RLHF), it is possible to fine-tune unsupervised LMs to follow instructions and produce outputs that reflect human preferences. Despite its benefits, RLHF has been shown to potentially harm a language model's reasoning capabilities and introduce artifacts such as hallucinations where the model may fabricate facts. To address this issue we introduce Direct Preference Heads (DPH), a fine-tuning framework that enables LMs to learn human preference signals through an auxiliary reward head without directly affecting the output distribution of the language modeling head. We perform a theoretical analysis of our objective function and find strong ties to Conservative Direct Preference Optimization (cDPO). Finally we evaluate our models on GLUE, RACE, and the GPT4All evaluation suite and demonstrate that our method produces models which achieve higher scores than those fine-tuned with Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO) alone.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20053
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads
Hadji-Kyriacou, Avelina Asada
Arandjelovic, Ognjen
Computation and Language
Artificial Intelligence
Machine Learning
Pre-trained Language Models (LMs) exhibit strong zero-shot and in-context learning capabilities; however, their behaviors are often difficult to control. By utilizing Reinforcement Learning from Human Feedback (RLHF), it is possible to fine-tune unsupervised LMs to follow instructions and produce outputs that reflect human preferences. Despite its benefits, RLHF has been shown to potentially harm a language model's reasoning capabilities and introduce artifacts such as hallucinations where the model may fabricate facts. To address this issue we introduce Direct Preference Heads (DPH), a fine-tuning framework that enables LMs to learn human preference signals through an auxiliary reward head without directly affecting the output distribution of the language modeling head. We perform a theoretical analysis of our objective function and find strong ties to Conservative Direct Preference Optimization (cDPO). Finally we evaluate our models on GLUE, RACE, and the GPT4All evaluation suite and demonstrate that our method produces models which achieve higher scores than those fine-tuned with Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO) alone.
title Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.20053