Semi-Supervised Preference Optimization with Limited Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Seonggyun, Lim, Sungjun, Park, Seojin, Cheon, Soeun, Song, Kyungwoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911456357777408
author Lee, Seonggyun
Lim, Sungjun
Park, Seojin
Cheon, Soeun
Song, Kyungwoo
author_facet Lee, Seonggyun
Lim, Sungjun
Park, Seojin
Cheon, Soeun
Song, Kyungwoo
contents The field of preference optimization has made outstanding contributions to the alignment of language models with human preferences. Despite these advancements, recent methods still rely heavily on substantial paired (labeled) feedback data, leading to substantial resource expenditures. To address these challenges, we study the problem of Semi-Supervised Preference Optimization (SSPO) in which the idea is to learn from both a small number of pairwise preference labels and a large pool of unpaired samples simultaneously. Our key theoretical contribution proves the existence of an optimal reward threshold capable of separating winning and losing responses with high probability, which enables a principled pseudo-labeling of unpaired data. By leveraging these pseudo-labels, SSPO effectively distills latent preferences from large-scale unpaired data, thus maintaining human alignment while drastically reducing acquisition costs. Extensive experiments across datasets validate this remarkable data efficiency; for instance, SSPO trained with Mistral-7B-Instruct on just 1% of UltraFeedback consistently surpasses strong baselines trained on 10% of UltraFeedback.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semi-Supervised Preference Optimization with Limited Feedback
Lee, Seonggyun
Lim, Sungjun
Park, Seojin
Cheon, Soeun
Song, Kyungwoo
Machine Learning
Artificial Intelligence
The field of preference optimization has made outstanding contributions to the alignment of language models with human preferences. Despite these advancements, recent methods still rely heavily on substantial paired (labeled) feedback data, leading to substantial resource expenditures. To address these challenges, we study the problem of Semi-Supervised Preference Optimization (SSPO) in which the idea is to learn from both a small number of pairwise preference labels and a large pool of unpaired samples simultaneously. Our key theoretical contribution proves the existence of an optimal reward threshold capable of separating winning and losing responses with high probability, which enables a principled pseudo-labeling of unpaired data. By leveraging these pseudo-labels, SSPO effectively distills latent preferences from large-scale unpaired data, thus maintaining human alignment while drastically reducing acquisition costs. Extensive experiments across datasets validate this remarkable data efficiency; for instance, SSPO trained with Mistral-7B-Instruct on just 1% of UltraFeedback consistently surpasses strong baselines trained on 10% of UltraFeedback.
title Semi-Supervised Preference Optimization with Limited Feedback
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.00040