MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Vári-Kakas, Andor, Park, Ji Won, Tagasovska, Natasa |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BOtied: Multi-objective Bayesian optimization with tied multivariate ranks
by: Park, Ji Won, et al.
Published: (2023)
by: Park, Ji Won, et al.
Published: (2023)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
Implicitly Guided Design with PropEn: Match your Data to Follow the Gradient
by: Tagasovska, Nataša, et al.
Published: (2024)
by: Tagasovska, Nataša, et al.
Published: (2024)
MGDA Converges under Generalized Smoothness, Provably
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Finding Colorings in One-Sided Expanders
by: Buhai, Rares-Darius, et al.
Published: (2025)
by: Buhai, Rares-Darius, et al.
Published: (2025)
MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning
by: Lei, Xing, et al.
Published: (2024)
by: Lei, Xing, et al.
Published: (2024)
Provably Convergent Primal-Dual DPO for Constrained LLM Alignment
by: Du, Yihan, et al.
Published: (2025)
by: Du, Yihan, et al.
Published: (2025)
Antibody DomainBed: Out-of-Distribution Generalization in Therapeutic Protein Design
by: Tagasovska, Nataša, et al.
Published: (2024)
by: Tagasovska, Nataša, et al.
Published: (2024)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
by: Pal, Arka, et al.
Published: (2024)
by: Pal, Arka, et al.
Published: (2024)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
by: Ren, Yinuo, et al.
Published: (2024)
by: Ren, Yinuo, et al.
Published: (2024)
Supervised Contrastive Block Disentanglement
by: Makino, Taro, et al.
Published: (2025)
by: Makino, Taro, et al.
Published: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025)
by: Zhang, Zhengze, et al.
Published: (2025)
Uncovering Cross-Objective Interference in Multi-Objective Alignment
by: Lu, Yining, et al.
Published: (2026)
by: Lu, Yining, et al.
Published: (2026)
The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
by: Pollanen, Marco
Published: (2026)
by: Pollanen, Marco
Published: (2026)
VERI-DPO: Evidence-Aware Alignment for Clinical Summarization via Claim Verification and Direct Preference Optimization
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
Critical Patch-Aware Sparse Prompting with Decoupled Training for Continual Learning on the Edge
by: Lim, Wonseon, et al.
Published: (2026)
by: Lim, Wonseon, et al.
Published: (2026)
Uncertainty modeling for fine-tuned implicit functions
by: Susmelj, Anna, et al.
Published: (2024)
by: Susmelj, Anna, et al.
Published: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
Knowledge Gradient for Multi-Objective Bayesian Optimization with Decoupled Evaluations
by: Buckingham, Jack M., et al.
Published: (2023)
by: Buckingham, Jack M., et al.
Published: (2023)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
by: Liu, Jilong, et al.
Published: (2026)
by: Liu, Jilong, et al.
Published: (2026)
MIRACL: A Diverse Meta-Reinforcement Learning for Multi-Objective Multi-Echelon Combinatorial Supply Chain Optimisation
by: Rachman, Rifny, et al.
Published: (2026)
by: Rachman, Rifny, et al.
Published: (2026)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
by: Lin, Baijiong, et al.
Published: (2025)
by: Lin, Baijiong, et al.
Published: (2025)
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
by: Deng, Yihe, et al.
Published: (2024)
by: Deng, Yihe, et al.
Published: (2024)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
A Decoupled Basis-Vector-Driven Generative Framework for Dynamic Multi-Objective Optimization
by: Yang, Yaoming, et al.
Published: (2026)
by: Yang, Yaoming, et al.
Published: (2026)
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
by: Beikzadeh, Mehrab, et al.
Published: (2026)
by: Beikzadeh, Mehrab, et al.
Published: (2026)
Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
by: Li, Yining, et al.
Published: (2026)
by: Li, Yining, et al.
Published: (2026)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Center-Outward q-Dominance: A Sample-Computable Proxy for Strong Stochastic Dominance in Multi-Objective Optimisation
by: van der Laag, Robin, et al.
Published: (2025)
by: van der Laag, Robin, et al.
Published: (2025)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
by: Yang, Zhiqin, et al.
Published: (2026)
by: Yang, Zhiqin, et al.
Published: (2026)
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
by: Yang, Yonghui, et al.
Published: (2026)
by: Yang, Yonghui, et al.
Published: (2026)
Decoupled Conformal Optimisation: Efficient Prediction Sets via Independent Tuning and Calibration
by: Wu, Fanyi, et al.
Published: (2026)
by: Wu, Fanyi, et al.
Published: (2026)
Reward Dimension Reduction for Scalable Multi-Objective Reinforcement Learning
by: Park, Giseung, et al.
Published: (2025)
by: Park, Giseung, et al.
Published: (2025)
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
by: Bou, Matthieu, et al.
Published: (2025)
by: Bou, Matthieu, et al.
Published: (2025)
Improving Discrete Optimisation Via Decoupled Straight-Through Estimator
by: Shah, Rushi, et al.
Published: (2024)
by: Shah, Rushi, et al.
Published: (2024)
Similar Items
-
BOtied: Multi-objective Bayesian optimization with tied multivariate ranks
by: Park, Ji Won, et al.
Published: (2023) -
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025) -
Implicitly Guided Design with PropEn: Match your Data to Follow the Gradient
by: Tagasovska, Nataša, et al.
Published: (2024) -
MGDA Converges under Generalized Smoothness, Provably
by: Zhang, Qi, et al.
Published: (2024) -
Finding Colorings in One-Sided Expanders
by: Buhai, Rares-Darius, et al.
Published: (2025)