Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Sen, Cui, Leyang, Cai, Deng, Huang, Xinting, Shi, Shuming, Lam, Wai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914969832914944
author Yang, Sen
Cui, Leyang
Cai, Deng
Huang, Xinting
Shi, Shuming
Lam, Wai
author_facet Yang, Sen
Cui, Leyang
Cai, Deng
Huang, Xinting
Shi, Shuming
Lam, Wai
contents Iterative preference learning, though yielding superior performances, requires online annotated preference labels. In this work, we study strategies to select worth-annotating response pairs for cost-efficient annotation while achieving competitive or even better performances compared with the random selection baseline for iterative preference learning. Built on assumptions regarding uncertainty and distribution shifts, we propose a comparative view to rank the implicit reward margins as predicted by DPO to select the response pairs that yield more benefits. Through extensive experiments, we show that annotating those response pairs with small margins is generally better than large or random, under both single- and multi-iteration scenarios. Besides, our empirical results suggest allocating more annotation budgets in the earlier iterations rather than later across multiple iterations.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17312
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
Yang, Sen
Cui, Leyang
Cai, Deng
Huang, Xinting
Shi, Shuming
Lam, Wai
Computation and Language
Iterative preference learning, though yielding superior performances, requires online annotated preference labels. In this work, we study strategies to select worth-annotating response pairs for cost-efficient annotation while achieving competitive or even better performances compared with the random selection baseline for iterative preference learning. Built on assumptions regarding uncertainty and distribution shifts, we propose a comparative view to rank the implicit reward margins as predicted by DPO to select the response pairs that yield more benefits. Through extensive experiments, we show that annotating those response pairs with small margins is generally better than large or random, under both single- and multi-iteration scenarios. Besides, our empirical results suggest allocating more annotation budgets in the earlier iterations rather than later across multiple iterations.
title Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
topic Computation and Language
url https://arxiv.org/abs/2406.17312