Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Yuxiao, Sinha, Arunesh, Varakantham, Pradeep
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929623903764480
author Lu, Yuxiao
Sinha, Arunesh
Varakantham, Pradeep
author_facet Lu, Yuxiao
Sinha, Arunesh
Varakantham, Pradeep
contents Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, previous approaches often demand substantial human data collection or rely on the less dependable option of using another LLM to generate corrective data. In this paper, we aim to take this problem and overcome limitations of requiring significant high-quality human data. Our method requires only a small set of unsafe responses to toxic prompts, easily obtained from the unsafe LLM itself. By employing a semantic cost combined with a negative Earth Mover Distance (EMD) loss, we guide the LLM away from generating unsafe responses. Additionally, we propose a novel lower bound for EMD loss, enabling more efficient optimization. Our results demonstrate superior performance and data efficiency compared to baselines, and we further examine the nuanced effects of over-alignment and potential degradation of language capabilities when using contrastive data.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06843
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
Lu, Yuxiao
Sinha, Arunesh
Varakantham, Pradeep
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, previous approaches often demand substantial human data collection or rely on the less dependable option of using another LLM to generate corrective data. In this paper, we aim to take this problem and overcome limitations of requiring significant high-quality human data. Our method requires only a small set of unsafe responses to toxic prompts, easily obtained from the unsafe LLM itself. By employing a semantic cost combined with a negative Earth Mover Distance (EMD) loss, we guide the LLM away from generating unsafe responses. Additionally, we propose a novel lower bound for EMD loss, enabling more efficient optimization. Our results demonstrate superior performance and data efficiency compared to baselines, and we further examine the nuanced effects of over-alignment and potential degradation of language capabilities when using contrastive data.
title Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.06843