Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909466352418816 |
|---|---|
| author | Diwan, Nirav Ergen, Tolga Shim, Dongsub Lee, Honglak |
| author_facet | Diwan, Nirav Ergen, Tolga Shim, Dongsub Lee, Honglak |
| contents | Direct Preference Optimization (DPO) has emerged as a de-facto approach for aligning language models with human preferences. Recent work has shown DPO's effectiveness relies on training data quality. In particular, clear quality differences between preferred and rejected responses enhance learning performance. Current methods for identifying and obtaining such high-quality samples demand additional resources or external models. We discover that reference model probability space naturally detects high-quality training samples. Using this insight, we present a sampling strategy that achieves consistent improvements (+0.1 to +0.4) on MT-Bench while using less than half (30-50%) of the training data. We observe substantial improvements (+0.4 to +0.98) for technical tasks (coding, math, and reasoning) across multiple models and hyperparameter settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_15109 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Diwan, Nirav Ergen, Tolga Shim, Dongsub Lee, Honglak Machine Learning Artificial Intelligence Direct Preference Optimization (DPO) has emerged as a de-facto approach for aligning language models with human preferences. Recent work has shown DPO's effectiveness relies on training data quality. In particular, clear quality differences between preferred and rejected responses enhance learning performance. Current methods for identifying and obtaining such high-quality samples demand additional resources or external models. We discover that reference model probability space naturally detects high-quality training samples. Using this insight, we present a sampling strategy that achieves consistent improvements (+0.1 to +0.4) on MT-Bench while using less than half (30-50%) of the training data. We observe substantial improvements (+0.4 to +0.98) for technical tasks (coding, math, and reasoning) across multiple models and hyperparameter settings. |
| title | Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2501.15109 |