Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914969832914944 |
|---|---|
| author | Yang, Sen Cui, Leyang Cai, Deng Huang, Xinting Shi, Shuming Lam, Wai |
| author_facet | Yang, Sen Cui, Leyang Cai, Deng Huang, Xinting Shi, Shuming Lam, Wai |
| contents | Iterative preference learning, though yielding superior performances, requires online annotated preference labels. In this work, we study strategies to select worth-annotating response pairs for cost-efficient annotation while achieving competitive or even better performances compared with the random selection baseline for iterative preference learning. Built on assumptions regarding uncertainty and distribution shifts, we propose a comparative view to rank the implicit reward margins as predicted by DPO to select the response pairs that yield more benefits. Through extensive experiments, we show that annotating those response pairs with small margins is generally better than large or random, under both single- and multi-iteration scenarios. Besides, our empirical results suggest allocating more annotation budgets in the earlier iterations rather than later across multiple iterations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_17312 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning Yang, Sen Cui, Leyang Cai, Deng Huang, Xinting Shi, Shuming Lam, Wai Computation and Language Iterative preference learning, though yielding superior performances, requires online annotated preference labels. In this work, we study strategies to select worth-annotating response pairs for cost-efficient annotation while achieving competitive or even better performances compared with the random selection baseline for iterative preference learning. Built on assumptions regarding uncertainty and distribution shifts, we propose a comparative view to rank the implicit reward margins as predicted by DPO to select the response pairs that yield more benefits. Through extensive experiments, we show that annotating those response pairs with small margins is generally better than large or random, under both single- and multi-iteration scenarios. Besides, our empirical results suggest allocating more annotation budgets in the earlier iterations rather than later across multiple iterations. |
| title | Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2406.17312 |