Saved in:
Bibliographic Details
Main Authors: Tuan, Yi-Lin, Wang, William Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2408.16751
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929478665502720
author Tuan, Yi-Lin
Wang, William Yang
author_facet Tuan, Yi-Lin
Wang, William Yang
contents Beyond maximum likelihood estimation (MLE), the standard objective of a language model (LM) that optimizes good examples probabilities, many studies have explored ways that also penalize bad examples for enhancing the quality of output distribution, including unlikelihood training, exponential maximizing average treatment effect (ExMATE), and direct preference optimization (DPO). To systematically compare these methods and further provide a unified recipe for LM optimization, in this paper, we present a unique angle of gradient analysis of loss functions that simultaneously reward good examples and penalize bad ones in LMs. Through both mathematical results and experiments on CausalDialogue and Anthropic HH-RLHF datasets, we identify distinct functional characteristics among these methods. We find that ExMATE serves as a superior surrogate for MLE, and that combining DPO with ExMATE instead of MLE further enhances both the statistical (5-7%) and generative (+18% win rate) performance.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models
Tuan, Yi-Lin
Wang, William Yang
Computation and Language
Machine Learning
Beyond maximum likelihood estimation (MLE), the standard objective of a language model (LM) that optimizes good examples probabilities, many studies have explored ways that also penalize bad examples for enhancing the quality of output distribution, including unlikelihood training, exponential maximizing average treatment effect (ExMATE), and direct preference optimization (DPO). To systematically compare these methods and further provide a unified recipe for LM optimization, in this paper, we present a unique angle of gradient analysis of loss functions that simultaneously reward good examples and penalize bad ones in LMs. Through both mathematical results and experiments on CausalDialogue and Anthropic HH-RLHF datasets, we identify distinct functional characteristics among these methods. We find that ExMATE serves as a superior surrogate for MLE, and that combining DPO with ExMATE instead of MLE further enhances both the statistical (5-7%) and generative (+18% win rate) performance.
title A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2408.16751