Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mukherjee, Subhojyoti, Lalitha, Anusha, Sengupta, Sailik, Deshmukh, Aniket, Kveton, Branislav
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913600603422720
author Mukherjee, Subhojyoti
Lalitha, Anusha
Sengupta, Sailik
Deshmukh, Aniket
Kveton, Branislav
author_facet Mukherjee, Subhojyoti
Lalitha, Anusha
Sengupta, Sailik
Deshmukh, Aniket
Kveton, Branislav
contents Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting. Recent works on MOAHF considered a-priori multi-objective optimization (MOO), where human preferences are known at training or inference time. In contrast, when human preferences are unknown or difficult to quantify, a natural approach is to cover the Pareto front by multiple diverse solutions. We propose an algorithm HaM for learning diverse LLM policies that maximizes their hypervolume. This is the first application of a-posteriori MOO to MOAHF. HaM is computationally and space efficient, and empirically superior across objectives such as harmlessness, helpfulness, humor, faithfulness, and hallucination, on various datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05469
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
Mukherjee, Subhojyoti
Lalitha, Anusha
Sengupta, Sailik
Deshmukh, Aniket
Kveton, Branislav
Machine Learning
Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting. Recent works on MOAHF considered a-priori multi-objective optimization (MOO), where human preferences are known at training or inference time. In contrast, when human preferences are unknown or difficult to quantify, a natural approach is to cover the Pareto front by multiple diverse solutions. We propose an algorithm HaM for learning diverse LLM policies that maximizes their hypervolume. This is the first application of a-posteriori MOO to MOAHF. HaM is computationally and space efficient, and empirically superior across objectives such as harmlessness, helpfulness, humor, faithfulness, and hallucination, on various datasets.
title Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
topic Machine Learning
url https://arxiv.org/abs/2412.05469