Meta-Learning Objectives for Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alfano, Carlo, Sapora, Silvia, Foerster, Jakob Nicolaus, Rebeschini, Patrick, Teh, Yee Whye
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909983893880832
author Alfano, Carlo
Sapora, Silvia
Foerster, Jakob Nicolaus
Rebeschini, Patrick
Teh, Yee Whye
author_facet Alfano, Carlo
Sapora, Silvia
Foerster, Jakob Nicolaus
Rebeschini, Patrick
Teh, Yee Whye
contents Evaluating preference optimization (PO) algorithms on LLM alignment is a challenging task that presents prohibitive costs, noise, and several variables like model size and hyper-parameters. In this work, we show that it is possible to gain insights on the efficacy of PO algorithm on simpler benchmarks. We design a diagnostic suite of MuJoCo tasks and datasets, which we use to systematically evaluate PO algorithms, establishing a more controlled and cheaper benchmark. We then propose a novel family of PO algorithms based on mirror descent, which we call Mirror Preference Optimization (MPO). Through evolutionary strategies, we search this class to discover algorithms specialized to specific properties of preference datasets, such as mixed-quality or noisy data. We demonstrate that our discovered PO algorithms outperform all known algorithms in the targeted MuJoCo settings. Finally, based on the insights gained from our MuJoCo experiments, we design a PO algorithm that significantly outperform existing baselines in an LLM alignment task.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06568
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Meta-Learning Objectives for Preference Optimization
Alfano, Carlo
Sapora, Silvia
Foerster, Jakob Nicolaus
Rebeschini, Patrick
Teh, Yee Whye
Machine Learning
Artificial Intelligence
Evaluating preference optimization (PO) algorithms on LLM alignment is a challenging task that presents prohibitive costs, noise, and several variables like model size and hyper-parameters. In this work, we show that it is possible to gain insights on the efficacy of PO algorithm on simpler benchmarks. We design a diagnostic suite of MuJoCo tasks and datasets, which we use to systematically evaluate PO algorithms, establishing a more controlled and cheaper benchmark. We then propose a novel family of PO algorithms based on mirror descent, which we call Mirror Preference Optimization (MPO). Through evolutionary strategies, we search this class to discover algorithms specialized to specific properties of preference datasets, such as mixed-quality or noisy data. We demonstrate that our discovered PO algorithms outperform all known algorithms in the targeted MuJoCo settings. Finally, based on the insights gained from our MuJoCo experiments, we design a PO algorithm that significantly outperform existing baselines in an LLM alignment task.
title Meta-Learning Objectives for Preference Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2411.06568