Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Wenxuan, Torr, Philip H. S., Elhoseiny, Mohamed, Bibi, Adel
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909570595553280
author Zhang, Wenxuan
Torr, Philip H. S.
Elhoseiny, Mohamed
Bibi, Adel
author_facet Zhang, Wenxuan
Torr, Philip H. S.
Elhoseiny, Mohamed
Bibi, Adel
contents Fine-tuning large language models (LLMs) on human preferences, typically through reinforcement learning from human feedback (RLHF), has proven successful in enhancing their capabilities. However, ensuring the safety of LLMs during fine-tuning remains a critical concern, and mitigating the potential conflicts in safety and helpfulness is costly in RLHF. To address this issue, we propose a supervised learning framework called Bi-Factorial Preference Optimization (BFPO), which re-parameterizes a joint RLHF objective of both safety and helpfulness into a single supervised learning objective. In supervised optimization, a labeling function is used to capture the global preferences ranking to balance both safety and helpfulness. To evaluate BFPO, we develop a benchmark that includes comprehensive discriminative and generative tasks for helpfulness and harmlessness. The results indicate that our method significantly outperforms existing approaches in both safety and helpfulness. Moreover, BFPO achieves the same level of safety as methods that heavily rely on human labor with less than 10\% of the computational resources and human prompting and annotation process. The training recipes can be found here: https://github.com/wx-zhang/bfpo.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
Zhang, Wenxuan
Torr, Philip H. S.
Elhoseiny, Mohamed
Bibi, Adel
Artificial Intelligence
Computation and Language
Machine Learning
Fine-tuning large language models (LLMs) on human preferences, typically through reinforcement learning from human feedback (RLHF), has proven successful in enhancing their capabilities. However, ensuring the safety of LLMs during fine-tuning remains a critical concern, and mitigating the potential conflicts in safety and helpfulness is costly in RLHF. To address this issue, we propose a supervised learning framework called Bi-Factorial Preference Optimization (BFPO), which re-parameterizes a joint RLHF objective of both safety and helpfulness into a single supervised learning objective. In supervised optimization, a labeling function is used to capture the global preferences ranking to balance both safety and helpfulness. To evaluate BFPO, we develop a benchmark that includes comprehensive discriminative and generative tasks for helpfulness and harmlessness. The results indicate that our method significantly outperforms existing approaches in both safety and helpfulness. Moreover, BFPO achieves the same level of safety as methods that heavily rely on human labor with less than 10\% of the computational resources and human prompting and annotation process. The training recipes can be found here: https://github.com/wx-zhang/bfpo.
title Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2408.15313