Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Yiran, Xia, Yu, Chang, Jonathan, Ammanabrolu, Prithviraj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913175265345536
author Shen, Yiran
Xia, Yu
Chang, Jonathan
Ammanabrolu, Prithviraj
author_facet Shen, Yiran
Xia, Yu
Chang, Jonathan
Ammanabrolu, Prithviraj
contents Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective. We seek to answer what it would take to simultaneously align a model across various domains spanning those with: verifiable rewards, non-verifiable subjective preferences, and complex interactive scenarios. Such multi-objective alignment setups are often plagued by individual objectives being at odds with each other, resulting in inefficient training and limited user control during inference. To address these issues, we propose $\textbf{M}$ulti-$\textbf{A}$ction-$\textbf{H}$ead $\textbf{AL}$ignment with PRM-guided Dec$\textbf{O}$ding ($\textbf{MAHALO}$), a unified framework that standardizes PRM training across verifiable and non-verifiable settings for step-level supervision, performs vectorized multi-objective alignment with Multi-Action-Head DPO, and enables controllable inference through objective-specific weighting and PRM-guided decoding. Experiments across math reasoning, human values alignment, and multi-turn tutoring show that MAHALO jointly improves multiple objectives simultaneously with limited interference, while remaining generalizable and adaptable across domains and offering flexible user control at inference time. Our code is available at: https://github.com/pearls-lab/multiobj-align.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
Shen, Yiran
Xia, Yu
Chang, Jonathan
Ammanabrolu, Prithviraj
Machine Learning
Artificial Intelligence
Computation and Language
Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective. We seek to answer what it would take to simultaneously align a model across various domains spanning those with: verifiable rewards, non-verifiable subjective preferences, and complex interactive scenarios. Such multi-objective alignment setups are often plagued by individual objectives being at odds with each other, resulting in inefficient training and limited user control during inference. To address these issues, we propose $\textbf{M}$ulti-$\textbf{A}$ction-$\textbf{H}$ead $\textbf{AL}$ignment with PRM-guided Dec$\textbf{O}$ding ($\textbf{MAHALO}$), a unified framework that standardizes PRM training across verifiable and non-verifiable settings for step-level supervision, performs vectorized multi-objective alignment with Multi-Action-Head DPO, and enables controllable inference through objective-specific weighting and PRM-guided decoding. Experiments across math reasoning, human values alignment, and multi-turn tutoring show that MAHALO jointly improves multiple objectives simultaneously with limited interference, while remaining generalizable and adaptable across domains and offering flexible user control at inference time. Our code is available at: https://github.com/pearls-lab/multiobj-align.
title Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.01167