Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Tong, Mei, Jincheng, Dai, Hanjun, Wen, Zixin, Cen, Shicong, Schuurmans, Dale, Chi, Yuejie, Dai, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915159872634880
author Yang, Tong
Mei, Jincheng
Dai, Hanjun
Wen, Zixin
Cen, Shicong
Schuurmans, Dale
Chi, Yuejie
Dai, Bo
author_facet Yang, Tong
Mei, Jincheng
Dai, Hanjun
Wen, Zixin
Cen, Shicong
Schuurmans, Dale
Chi, Yuejie
Dai, Bo
contents Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algorithm is prohibitively expensive in practice due to the sample and computation inefficiency. This paper addresses the problem by revealing a unified game-theoretic connection between iterative BOND and self-play alignment, which unifies seemingly disparate algorithmic paradigms. Based on the connection, we establish a novel framework, WIN rate Dominance (WIND), with a series of efficient algorithms for regularized win rate dominance optimization that approximates iterative BOND in the parameter space. We provides provable sample efficiency guarantee for one of the WIND variant with the square loss objective. The experimental results confirm that our algorithm not only accelerates the computation, but also achieves superior sample efficiency compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20727
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
Yang, Tong
Mei, Jincheng
Dai, Hanjun
Wen, Zixin
Cen, Shicong
Schuurmans, Dale
Chi, Yuejie
Dai, Bo
Machine Learning
Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algorithm is prohibitively expensive in practice due to the sample and computation inefficiency. This paper addresses the problem by revealing a unified game-theoretic connection between iterative BOND and self-play alignment, which unifies seemingly disparate algorithmic paradigms. Based on the connection, we establish a novel framework, WIN rate Dominance (WIND), with a series of efficient algorithms for regularized win rate dominance optimization that approximates iterative BOND in the parameter space. We provides provable sample efficiency guarantee for one of the WIND variant with the square loss objective. The experimental results confirm that our algorithm not only accelerates the computation, but also achieves superior sample efficiency compared to existing methods.
title Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
topic Machine Learning
url https://arxiv.org/abs/2410.20727