When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Afzali, Amirabbas, Jeon, Myeongho, Brbic, Maria
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917315982917632
author Afzali, Amirabbas
Jeon, Myeongho
Brbic, Maria
author_facet Afzali, Amirabbas
Jeon, Myeongho
Brbic, Maria
contents Preference alignment is an essential step in adapting large language models (LLMs) to human values, but existing approaches typically depend on costly human annotations or large-scale API-based models. We explore whether a weak LLM can instead act as an effective annotator. We surprisingly find that selecting only a subset of a weak LLM's highly confident samples leads to substantially better performance than using full human annotations. Building on this insight, we propose Confidence-Weighted Preference Optimization (CW-PO), a general framework that re-weights training samples by a weak LLM's confidence and can be applied across different preference optimization objectives. Notably, the model aligned by CW-PO with just 20% of human annotations outperforms the model trained with 100% of annotations under standard DPO. These results suggest that weak LLMs, when paired with confidence weighting, can dramatically reduce the cost of preference alignment while even outperforming methods trained on fully human-labeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04968
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
Afzali, Amirabbas
Jeon, Myeongho
Brbic, Maria
Computation and Language
Artificial Intelligence
Preference alignment is an essential step in adapting large language models (LLMs) to human values, but existing approaches typically depend on costly human annotations or large-scale API-based models. We explore whether a weak LLM can instead act as an effective annotator. We surprisingly find that selecting only a subset of a weak LLM's highly confident samples leads to substantially better performance than using full human annotations. Building on this insight, we propose Confidence-Weighted Preference Optimization (CW-PO), a general framework that re-weights training samples by a weak LLM's confidence and can be applied across different preference optimization objectives. Notably, the model aligned by CW-PO with just 20% of human annotations outperforms the model trained with 100% of annotations under standard DPO. These results suggest that weak LLMs, when paired with confidence weighting, can dramatically reduce the cost of preference alignment while even outperforming methods trained on fully human-labeled data.
title When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.04968