Saved in:
Bibliographic Details
Main Authors: Jain, Samyak, Lubana, Ekdeep Singh, Oksuz, Kemal, Joy, Tom, Torr, Philip H. S., Sanyal, Amartya, Dokania, Puneet K.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.10264
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913475087826944
author Jain, Samyak
Lubana, Ekdeep Singh
Oksuz, Kemal
Joy, Tom
Torr, Philip H. S.
Sanyal, Amartya
Dokania, Puneet K.
author_facet Jain, Samyak
Lubana, Ekdeep Singh
Oksuz, Kemal
Joy, Tom
Torr, Philip H. S.
Sanyal, Amartya
Dokania, Puneet K.
contents Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via safety fine-tuning, we design a synthetic data generation framework that captures salient aspects of an unsafe input by modeling the interaction between the task the model is asked to perform (e.g., "design") versus the specific concepts the task is asked to be performed upon (e.g., a "cycle" vs. a "bomb"). Using this, we investigate three well-known safety fine-tuning methods -- supervised safety fine-tuning, direct preference optimization, and unlearning -- and provide significant evidence demonstrating that these methods minimally transform MLP weights to specifically align unsafe inputs into its weights' null space. This yields a clustering of inputs based on whether the model deems them safe or not. Correspondingly, when an adversarial input (e.g., a jailbreak) is provided, its activations are closer to safer samples, leading to the model processing such an input as if it were safe. We validate our findings, wherever possible, on real-world models -- specifically, Llama-2 7B and Llama-3 8B.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10264
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
Jain, Samyak
Lubana, Ekdeep Singh
Oksuz, Kemal
Joy, Tom
Torr, Philip H. S.
Sanyal, Amartya
Dokania, Puneet K.
Machine Learning
Computation and Language
Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via safety fine-tuning, we design a synthetic data generation framework that captures salient aspects of an unsafe input by modeling the interaction between the task the model is asked to perform (e.g., "design") versus the specific concepts the task is asked to be performed upon (e.g., a "cycle" vs. a "bomb"). Using this, we investigate three well-known safety fine-tuning methods -- supervised safety fine-tuning, direct preference optimization, and unlearning -- and provide significant evidence demonstrating that these methods minimally transform MLP weights to specifically align unsafe inputs into its weights' null space. This yields a clustering of inputs based on whether the model deems them safe or not. Correspondingly, when an adversarial input (e.g., a jailbreak) is provided, its activations are closer to safer samples, leading to the model processing such an input as if it were safe. We validate our findings, wherever possible, on real-world models -- specifically, Llama-2 7B and Llama-3 8B.
title What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2407.10264