Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Feifan, Wei, Shaohang, Luo, Wen, Fan, Yuxuan, Liu, Tianyu, Wang, Guoyin, Wang, Houfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910995596705792
author Song, Feifan
Wei, Shaohang
Luo, Wen
Fan, Yuxuan
Liu, Tianyu
Wang, Guoyin
Wang, Houfeng
author_facet Song, Feifan
Wei, Shaohang
Luo, Wen
Fan, Yuxuan
Liu, Tianyu
Wang, Guoyin
Wang, Houfeng
contents Large Language Models (LLMs) require alignment with human preferences to avoid generating offensive, false, or meaningless content. Recently, low-resource methods for LLM alignment have been popular, while still facing challenges in obtaining both high-quality and aligned content. Motivated by the observation that the difficulty of generating aligned responses is concentrated at the beginning of decoding, we propose a novel framework, Weak-to-Strong Decoding (WSD), to enhance the alignment ability of base models by the guidance of a small aligned model. The small model first drafts well-aligned beginnings, followed by the large base model to continue the rest, controlled by a well-designed auto-switch mechanism. We also collect a new dataset, GenerAlign, to fine-tune a small-sized Pilot-3B as the draft model, which effectively enhances different base models under the WSD framework to outperform all baseline methods, while avoiding degradation on downstream tasks, termed as the alignment tax. Extensive experiments are further conducted to examine the impact of different settings and time efficiency, as well as analyses on the intrinsic mechanisms of WSD in depth.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
Song, Feifan
Wei, Shaohang
Luo, Wen
Fan, Yuxuan
Liu, Tianyu
Wang, Guoyin
Wang, Houfeng
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) require alignment with human preferences to avoid generating offensive, false, or meaningless content. Recently, low-resource methods for LLM alignment have been popular, while still facing challenges in obtaining both high-quality and aligned content. Motivated by the observation that the difficulty of generating aligned responses is concentrated at the beginning of decoding, we propose a novel framework, Weak-to-Strong Decoding (WSD), to enhance the alignment ability of base models by the guidance of a small aligned model. The small model first drafts well-aligned beginnings, followed by the large base model to continue the rest, controlled by a well-designed auto-switch mechanism. We also collect a new dataset, GenerAlign, to fine-tune a small-sized Pilot-3B as the draft model, which effectively enhances different base models under the WSD framework to outperform all baseline methods, while avoiding degradation on downstream tasks, termed as the alignment tax. Extensive experiments are further conducted to examine the impact of different settings and time efficiency, as well as analyses on the intrinsic mechanisms of WSD in depth.
title Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.07434