Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Junming, Xu, Ning, Liu, Biao, Qiao, Shiqi, Geng, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910034517032960
author Yang, Junming
Xu, Ning
Liu, Biao
Qiao, Shiqi
Geng, Xin
author_facet Yang, Junming
Xu, Ning
Liu, Biao
Qiao, Shiqi
Geng, Xin
contents Preference optimization is crucial for aligning large language models (LLMs) with human values and intentions. A significant challenge in this process is the distribution mismatch between pre-collected offline preference data and the evolving model policy. Existing methods attempt to reduce this gap using static heuristics or decoupled online sampling strategies, but they often fail to adapt to the model's dynamic learning state. To bridge this gap, we propose Meta-Weighted Adaptive Preference Optimization (MetaAPO), a novel framework that dynamically couples data generation with model training. MetaAPO employs a lightweight meta-learner, as an "alignment gap estimator", to evaluate the potential benefits of on-policy sampling in relation to offline data. This guides targeted online generation and assigns sample-wise meta-weights to the optimization objective, dynamically balancing the quality and distribution of online and offline data. Experiments on AlpacaEval 2, Arena-Hard and MT-Bench demonstrate that MetaAPO consistently outperforms existing preference optimization approaches across various settings, while reducing 42% in online annotation costs. Code is available at https://github.com/junming-yang/MetaAPO.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23371
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
Yang, Junming
Xu, Ning
Liu, Biao
Qiao, Shiqi
Geng, Xin
Computation and Language
Artificial Intelligence
Machine Learning
Preference optimization is crucial for aligning large language models (LLMs) with human values and intentions. A significant challenge in this process is the distribution mismatch between pre-collected offline preference data and the evolving model policy. Existing methods attempt to reduce this gap using static heuristics or decoupled online sampling strategies, but they often fail to adapt to the model's dynamic learning state. To bridge this gap, we propose Meta-Weighted Adaptive Preference Optimization (MetaAPO), a novel framework that dynamically couples data generation with model training. MetaAPO employs a lightweight meta-learner, as an "alignment gap estimator", to evaluate the potential benefits of on-policy sampling in relation to offline data. This guides targeted online generation and assigns sample-wise meta-weights to the optimization objective, dynamically balancing the quality and distribution of online and offline data. Experiments on AlpacaEval 2, Arena-Hard and MT-Bench demonstrate that MetaAPO consistently outperforms existing preference optimization approaches across various settings, while reducing 42% in online annotation costs. Code is available at https://github.com/junming-yang/MetaAPO.
title Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.23371