An Embarrassingly Simple Defense Against LLM Abliteration Attacks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shairah, Harethah Abu, Hammoud, Hasan Abed Al Kader, Ghanem, Bernard, Turkiyyah, George
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908578599665664
author Shairah, Harethah Abu
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
Turkiyyah, George
author_facet Shairah, Harethah Abu
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
Turkiyyah, George
contents Large language models (LLMs) are typically aligned to refuse harmful instructions through safety fine-tuning. A recent attack, termed abliteration, identifies and suppresses the single latent direction most responsible for refusal behavior, thereby enabling models to generate harmful content. We propose a defense that fundamentally alters how models express refusal. We construct an extended-refusal dataset in which responses to harmful prompts provide detailed justifications before refusing, distributing the refusal signal across multiple token positions. Fine-tuning Llama-2-7B-Chat and Qwen2.5-Instruct (1.5B and 3B parameters) on this dataset yields models that maintain high refusal rates under abliteration: refusal rates drop by at most 10%, compared to 70-80% drops in baseline models. Comprehensive evaluations of safety and utility demonstrate that extended-refusal fine-tuning effectively neutralizes abliteration attacks while preserving general model performance and enhancing robustness across multiple alignment scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19056
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Embarrassingly Simple Defense Against LLM Abliteration Attacks
Shairah, Harethah Abu
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
Turkiyyah, George
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) are typically aligned to refuse harmful instructions through safety fine-tuning. A recent attack, termed abliteration, identifies and suppresses the single latent direction most responsible for refusal behavior, thereby enabling models to generate harmful content. We propose a defense that fundamentally alters how models express refusal. We construct an extended-refusal dataset in which responses to harmful prompts provide detailed justifications before refusing, distributing the refusal signal across multiple token positions. Fine-tuning Llama-2-7B-Chat and Qwen2.5-Instruct (1.5B and 3B parameters) on this dataset yields models that maintain high refusal rates under abliteration: refusal rates drop by at most 10%, compared to 70-80% drops in baseline models. Comprehensive evaluations of safety and utility demonstrate that extended-refusal fine-tuning effectively neutralizes abliteration attacks while preserving general model performance and enhancing robustness across multiple alignment scenarios.
title An Embarrassingly Simple Defense Against LLM Abliteration Attacks
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.19056