Multi-Objective Reward and Preference Optimization: Theory and Algorithms
Fuente:
arXiv
Saved in:
| Main Author: | Agnihotri, Akhil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Best Policy Learning from Trajectory Preference Feedback
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
e-COP : Episodic Constrained Optimization of Policies
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Preference-Guided Diffusion for Multi-Objective Offline Optimization
by: Annadani, Yashas, et al.
Published: (2025)
by: Annadani, Yashas, et al.
Published: (2025)
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
by: Cao, Linfeng, et al.
Published: (2025)
by: Cao, Linfeng, et al.
Published: (2025)
User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization
by: Ip, Joshua Hang Sai, et al.
Published: (2025)
by: Ip, Joshua Hang Sai, et al.
Published: (2025)
Pareto Merging: Multi-Objective Optimization for Preference-Aware Model Merging
by: Chen, Weiyu, et al.
Published: (2024)
by: Chen, Weiyu, et al.
Published: (2024)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
by: Lin, Baijiong, et al.
Published: (2025)
by: Lin, Baijiong, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
Meta-Learning Objectives for Preference Optimization
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
Preference-based Multi-Objective Reinforcement Learning
by: Mu, Ni, et al.
Published: (2025)
by: Mu, Ni, et al.
Published: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Interactive Hyperparameter Optimization in Multi-Objective Problems via Preference Learning
by: Giovanelli, Joseph, et al.
Published: (2023)
by: Giovanelli, Joseph, et al.
Published: (2023)
Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond
by: Chen, Weiyu, et al.
Published: (2025)
by: Chen, Weiyu, et al.
Published: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Multiple Wasserstein Gradient Descent Algorithm for Multi-Objective Distributional Optimization
by: Nguyen, Dai Hai, et al.
Published: (2025)
by: Nguyen, Dai Hai, et al.
Published: (2025)
Explicit Preference Optimization: No Need for an Implicit Reward Model
by: Hu, Xiangkun, et al.
Published: (2025)
by: Hu, Xiangkun, et al.
Published: (2025)
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
by: Xu, Wenzhe, et al.
Published: (2026)
by: Xu, Wenzhe, et al.
Published: (2026)
A First-Order Multi-Gradient Algorithm for Multi-Objective Bi-Level Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Reward Dimension Reduction for Scalable Multi-Objective Reinforcement Learning
by: Park, Giseung, et al.
Published: (2025)
by: Park, Giseung, et al.
Published: (2025)
A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning
by: Chen, Ying-Tu, et al.
Published: (2026)
by: Chen, Ying-Tu, et al.
Published: (2026)
Mixture-Model Preference Learning for Many-Objective Bayesian Optimization
by: Dubey, Manisha, et al.
Published: (2026)
by: Dubey, Manisha, et al.
Published: (2026)
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
by: Ambadkar, Tanmay, et al.
Published: (2026)
by: Ambadkar, Tanmay, et al.
Published: (2026)
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
Best-of-Both-Worlds Multi-Dueling Bandits: Unified Algorithms for Stochastic and Adversarial Preferences under Condorcet and Borda Objectives
by: Akash, S, et al.
Published: (2026)
by: Akash, S, et al.
Published: (2026)
Learning Pareto-Optimal Rewards from Noisy Preferences: A Framework for Multi-Objective Inverse Reinforcement Learning
by: Cherukuri, Kalyan, et al.
Published: (2025)
by: Cherukuri, Kalyan, et al.
Published: (2025)
BOPO: Neural Combinatorial Optimization via Best-anchored and Objective-guided Preference Optimization
by: Liao, Zijun, et al.
Published: (2025)
by: Liao, Zijun, et al.
Published: (2025)
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
by: Shianifar, Jonaid, et al.
Published: (2026)
by: Shianifar, Jonaid, et al.
Published: (2026)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
by: Chen, Lisha, et al.
Published: (2025)
by: Chen, Lisha, et al.
Published: (2025)
Automatic Reward Shaping from Multi-Objective Human Heuristics
by: Xie, Yuqing, et al.
Published: (2025)
by: Xie, Yuqing, et al.
Published: (2025)
Similar Items
-
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025) -
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023) -
Best Policy Learning from Trajectory Preference Feedback
by: Agnihotri, Akhil, et al.
Published: (2025) -
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024) -
e-COP : Episodic Constrained Optimization of Policies
by: Agnihotri, Akhil, et al.
Published: (2024)