PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Dongxu, Xu, Bing, Chen, Yinzhuo, Xu, Bufan, Lu, Wenpeng, Yang, Muyun, Zhao, Tiejun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910681797754880
author Liu, Dongxu
Xu, Bing
Chen, Yinzhuo
Xu, Bufan
Lu, Wenpeng
Yang, Muyun
Zhao, Tiejun
author_facet Liu, Dongxu
Xu, Bing
Chen, Yinzhuo
Xu, Bufan
Lu, Wenpeng
Yang, Muyun
Zhao, Tiejun
contents Reinforcement Learning from Human Feedback (RLHF) has been proven to be an effective method for preference alignment of large language models (LLMs) and is widely used in the post-training process of LLMs. However, RLHF struggles with handling multiple competing preferences. This leads to a decrease in the alignment of LLMs with human preferences. To address this issue, we propose Preference Mixture of LoRAs (PMoL) from the perspective of model architecture, which can adapt to any number of preferences to mix. PMoL combines Mixture of Experts (MoE) and Low Rank Adaptor (LoRA). This architecture is innovatively applied to the research of preference alignment and has achieved significant performance improvement. The expert group soft loss is used to enable MoE with the ability to mix preferences. Through comprehensive evaluation by the reward model and GPT-4o, the experiment results show that PMoL has superior preference mixing capabilities compared to baseline methods. PMoL achieves better preference alignment with lower training costs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01245
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
Liu, Dongxu
Xu, Bing
Chen, Yinzhuo
Xu, Bufan
Lu, Wenpeng
Yang, Muyun
Zhao, Tiejun
Computation and Language
Reinforcement Learning from Human Feedback (RLHF) has been proven to be an effective method for preference alignment of large language models (LLMs) and is widely used in the post-training process of LLMs. However, RLHF struggles with handling multiple competing preferences. This leads to a decrease in the alignment of LLMs with human preferences. To address this issue, we propose Preference Mixture of LoRAs (PMoL) from the perspective of model architecture, which can adapt to any number of preferences to mix. PMoL combines Mixture of Experts (MoE) and Low Rank Adaptor (LoRA). This architecture is innovatively applied to the research of preference alignment and has achieved significant performance improvement. The expert group soft loss is used to enable MoE with the ability to mix preferences. Through comprehensive evaluation by the reward model and GPT-4o, the experiment results show that PMoL has superior preference mixing capabilities compared to baseline methods. PMoL achieves better preference alignment with lower training costs.
title PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
topic Computation and Language
url https://arxiv.org/abs/2411.01245