H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tekin, Selim Furkan, Ilhan, Fatih, Huang, Tiansheng, Hu, Sihao, Xu, Yichang, Yahn, Zachary, Liu, Ling
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917213001220096
author Tekin, Selim Furkan
Ilhan, Fatih
Huang, Tiansheng
Hu, Sihao
Xu, Yichang
Yahn, Zachary
Liu, Ling
author_facet Tekin, Selim Furkan
Ilhan, Fatih
Huang, Tiansheng
Hu, Sihao
Xu, Yichang
Yahn, Zachary
Liu, Ling
contents The alignment of pre-trained LLMs continues to draw significant attention from both industry and academia, aiming to ensure responses that are helpful, harmless, and honest. However, identifying a point in the model's representation subspace that simultaneously satisfies all these properties remains challenging. H3Fusion addresses this challenge by introducing a mixture-of-experts (MoE)-based fusion mechanism that models alignment as a controllable drift within the subspace, guided by a drift-regularization loss to balance competing alignment dimensions. Furthermore, we formulate the alignment by finding a dual objective of harnessing the distance of generated embeddings and alignment embeddings, and introduce a gating loss by canalizing the activations on the contributing experts. Extensive evaluations of three benchmark datasets show that H3Fusion is more helpful, less harmful, and more honest in three aspects: it outperforms each individually aligned model by 11.37%, and provides stronger robustness compared to the state-of-the-art LLM ensemble approaches by 13.77% and model-merging approaches by 6.18%. Code is available at https://github.com/git-disl/h3fusion.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17792
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
Tekin, Selim Furkan
Ilhan, Fatih
Huang, Tiansheng
Hu, Sihao
Xu, Yichang
Yahn, Zachary
Liu, Ling
Computation and Language
Artificial Intelligence
Machine Learning
The alignment of pre-trained LLMs continues to draw significant attention from both industry and academia, aiming to ensure responses that are helpful, harmless, and honest. However, identifying a point in the model's representation subspace that simultaneously satisfies all these properties remains challenging. H3Fusion addresses this challenge by introducing a mixture-of-experts (MoE)-based fusion mechanism that models alignment as a controllable drift within the subspace, guided by a drift-regularization loss to balance competing alignment dimensions. Furthermore, we formulate the alignment by finding a dual objective of harnessing the distance of generated embeddings and alignment embeddings, and introduce a gating loss by canalizing the activations on the contributing experts. Extensive evaluations of three benchmark datasets show that H3Fusion is more helpful, less harmful, and more honest in three aspects: it outperforms each individually aligned model by 11.37%, and provides stronger robustness compared to the state-of-the-art LLM ensemble approaches by 13.77% and model-merging approaches by 6.18%. Code is available at https://github.com/git-disl/h3fusion.
title H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.17792