Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Farn, Hua, Su, Hsuan, Kumar, Shachi H, Sahay, Saurav, Chen, Shang-Tse, Lee, Hung-yi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916922455490560
author Farn, Hua
Su, Hsuan
Kumar, Shachi H
Sahay, Saurav
Chen, Shang-Tse
Lee, Hung-yi
author_facet Farn, Hua
Su, Hsuan
Kumar, Shachi H
Sahay, Saurav
Chen, Shang-Tse
Lee, Hung-yi
contents Fine-tuning large language models (LLMs) for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of originally aligned models. While some existing methods attempt to restore safety by incorporating additional safety data, the quality of such data typically falls short of that used in the original alignment process. Moreover, these high-quality safety datasets are generally inaccessible, making it difficult to fully recover the model's original safety. We ask: How can we preserve safety while improving downstream task performance without additional safety data? We show that simply merging the weights of pre- and post-fine-tuned models effectively mitigates safety degradation while enhancing performance. Experiments across different downstream tasks and models validate the method's practicality and effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19512
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
Farn, Hua
Su, Hsuan
Kumar, Shachi H
Sahay, Saurav
Chen, Shang-Tse
Lee, Hung-yi
Computation and Language
Fine-tuning large language models (LLMs) for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of originally aligned models. While some existing methods attempt to restore safety by incorporating additional safety data, the quality of such data typically falls short of that used in the original alignment process. Moreover, these high-quality safety datasets are generally inaccessible, making it difficult to fully recover the model's original safety. We ask: How can we preserve safety while improving downstream task performance without additional safety data? We show that simply merging the weights of pre- and post-fine-tuned models effectively mitigates safety degradation while enhancing performance. Experiments across different downstream tasks and models validate the method's practicality and effectiveness.
title Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
topic Computation and Language
url https://arxiv.org/abs/2412.19512