DiffPO: Diffusion-styled Preference Optimization for Efficient Inference-Time Alignment of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ruizhe, Chai, Wenhao, Yang, Zhifei, Zhang, Xiaotian, Zhou, Joey Tianyi, Quek, Tony, Poria, Soujanya, Liu, Zuozhu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915302939295744
author Chen, Ruizhe
Chai, Wenhao
Yang, Zhifei
Zhang, Xiaotian
Zhou, Joey Tianyi
Quek, Tony
Poria, Soujanya
Liu, Zuozhu
author_facet Chen, Ruizhe
Chai, Wenhao
Yang, Zhifei
Zhang, Xiaotian
Zhou, Joey Tianyi
Quek, Tony
Poria, Soujanya
Liu, Zuozhu
contents Inference-time alignment provides an efficient alternative for aligning LLMs with humans. However, these approaches still face challenges, such as limited scalability due to policy-specific value functions and latency during the inference phase. In this paper, we propose a novel approach, Diffusion-styled Preference Optimization (\model), which provides an efficient and policy-agnostic solution for aligning LLMs with humans. By directly performing alignment at sentence level, \model~avoids the time latency associated with token-level generation. Designed as a plug-and-play module, \model~can be seamlessly integrated with various base models to enhance their alignment. Extensive experiments on AlpacaEval 2, MT-bench, and HH-RLHF demonstrate that \model~achieves superior alignment performance across various settings, achieving a favorable trade-off between alignment quality and inference-time latency. Furthermore, \model~demonstrates model-agnostic scalability, significantly improving the performance of large models such as Llama-3-70B.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffPO: Diffusion-styled Preference Optimization for Efficient Inference-Time Alignment of Large Language Models
Chen, Ruizhe
Chai, Wenhao
Yang, Zhifei
Zhang, Xiaotian
Zhou, Joey Tianyi
Quek, Tony
Poria, Soujanya
Liu, Zuozhu
Computation and Language
Inference-time alignment provides an efficient alternative for aligning LLMs with humans. However, these approaches still face challenges, such as limited scalability due to policy-specific value functions and latency during the inference phase. In this paper, we propose a novel approach, Diffusion-styled Preference Optimization (\model), which provides an efficient and policy-agnostic solution for aligning LLMs with humans. By directly performing alignment at sentence level, \model~avoids the time latency associated with token-level generation. Designed as a plug-and-play module, \model~can be seamlessly integrated with various base models to enhance their alignment. Extensive experiments on AlpacaEval 2, MT-bench, and HH-RLHF demonstrate that \model~achieves superior alignment performance across various settings, achieving a favorable trade-off between alignment quality and inference-time latency. Furthermore, \model~demonstrates model-agnostic scalability, significantly improving the performance of large models such as Llama-3-70B.
title DiffPO: Diffusion-styled Preference Optimization for Efficient Inference-Time Alignment of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2503.04240