Multi-objective Large Language Model Alignment with Hierarchical Experts
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908381324771328 |
|---|---|
| author | Li, Zhuo Du, Guodong Guo, Weiyang Zhou, Yigeng Li, Xiucheng Wang, Wenya Liu, Fangming Wang, Yequan Ye, Deheng Zhang, Min Li, Jing |
| author_facet | Li, Zhuo Du, Guodong Guo, Weiyang Zhou, Yigeng Li, Xiucheng Wang, Wenya Liu, Fangming Wang, Yequan Ye, Deheng Zhang, Min Li, Jing |
| contents | Aligning large language models (LLMs) to simultaneously satisfy multiple objectives remains a significant challenge, especially given the diverse and often conflicting nature of human preferences. Existing alignment methods struggle to balance trade-offs effectively, often requiring costly retraining or yielding suboptimal results across the Pareto frontier of preferences. In this paper, we introduce \textit{HoE}(Hierarchical Mixture-of-Experts), a \textit{lightweight}, \textit{parameter-efficient}, and \textit{plug-and-play} approach that eliminates the need for model training, while enabling LLMs to adapt across the entire Pareto frontier and accommodate diverse user preferences. In particular, \textit{HoE} consists of three hierarchical components: LoRA Experts, Router Experts and Preference Routing, reaching optimal Pareto frontiers and achieving a trade-off between parameter size, training cost, and performance. We evaluate \textit{HoE} across various tasks on 14 objectives and 200 different preferences among 6 benchmarks, demonstrating superior performance over 15 recent baselines. Code is available in the supplementary materials. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_20925 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multi-objective Large Language Model Alignment with Hierarchical Experts Li, Zhuo Du, Guodong Guo, Weiyang Zhou, Yigeng Li, Xiucheng Wang, Wenya Liu, Fangming Wang, Yequan Ye, Deheng Zhang, Min Li, Jing Computation and Language Artificial Intelligence Aligning large language models (LLMs) to simultaneously satisfy multiple objectives remains a significant challenge, especially given the diverse and often conflicting nature of human preferences. Existing alignment methods struggle to balance trade-offs effectively, often requiring costly retraining or yielding suboptimal results across the Pareto frontier of preferences. In this paper, we introduce \textit{HoE}(Hierarchical Mixture-of-Experts), a \textit{lightweight}, \textit{parameter-efficient}, and \textit{plug-and-play} approach that eliminates the need for model training, while enabling LLMs to adapt across the entire Pareto frontier and accommodate diverse user preferences. In particular, \textit{HoE} consists of three hierarchical components: LoRA Experts, Router Experts and Preference Routing, reaching optimal Pareto frontiers and achieving a trade-off between parameter size, training cost, and performance. We evaluate \textit{HoE} across various tasks on 14 objectives and 200 different preferences among 6 benchmarks, demonstrating superior performance over 15 recent baselines. Code is available in the supplementary materials. |
| title | Multi-objective Large Language Model Alignment with Hierarchical Experts |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.20925 |