Saved in:
Bibliographic Details
Main Authors: Liu, Xiao, Song, Xixuan, Dong, Yuxiao, Tang, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.00604
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929298563137536
author Liu, Xiao
Song, Xixuan
Dong, Yuxiao
Tang, Jie
author_facet Liu, Xiao
Song, Xixuan
Dong, Yuxiao
Tang, Jie
contents Reinforcement learning from human feedback (RLHF) has been a central technique for recent large language model (LLM) alignment. However, its heavy dependence on costly human or LLM-as-Judge preference feedback could stymie its wider applications. In this work, we introduce Self-Contrast, a feedback-free large language model alignment method via exploiting extensive self-generated negatives. With only supervised fine-tuning (SFT) targets, Self-Contrast leverages the LLM itself to generate massive diverse candidates, and harnesses a pre-trained embedding model to filter multiple negatives according to text similarity. Theoretically, we illustrate that in this setting, merely scaling negative responses can still effectively approximate situations with more balanced positive and negative preference annotations. Our experiments with direct preference optimization (DPO) on three datasets show that, Self-Contrast could consistently outperform SFT and standard DPO training by large margins. And as the number of self-generated negatives increases, the performance of Self-Contrast continues to grow. Code and data are available at https://github.com/THUDM/Self-Contrast.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00604
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
Liu, Xiao
Song, Xixuan
Dong, Yuxiao
Tang, Jie
Computation and Language
Artificial Intelligence
Machine Learning
Reinforcement learning from human feedback (RLHF) has been a central technique for recent large language model (LLM) alignment. However, its heavy dependence on costly human or LLM-as-Judge preference feedback could stymie its wider applications. In this work, we introduce Self-Contrast, a feedback-free large language model alignment method via exploiting extensive self-generated negatives. With only supervised fine-tuning (SFT) targets, Self-Contrast leverages the LLM itself to generate massive diverse candidates, and harnesses a pre-trained embedding model to filter multiple negatives according to text similarity. Theoretically, we illustrate that in this setting, merely scaling negative responses can still effectively approximate situations with more balanced positive and negative preference annotations. Our experiments with direct preference optimization (DPO) on three datasets show that, Self-Contrast could consistently outperform SFT and standard DPO training by large margins. And as the number of self-generated negatives increases, the performance of Self-Contrast continues to grow. Code and data are available at https://github.com/THUDM/Self-Contrast.
title Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.00604