RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asif, Sadia, Amiri, Mohammad Mohammadi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914527332794368
author Asif, Sadia
Amiri, Mohammad Mohammadi
author_facet Asif, Sadia
Amiri, Mohammad Mohammadi
contents Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment degradation. Our analysis shows that standard fine-tuning induces systematic drift in safety-relevant representations, distorts their geometric structure, and introduces interference between task optimization and safety features. These effects collectively lead to increased harmful compliance. Motivated by these findings, we introduce REFUSALGUARD, a representation-level fine-tuning framework that preserves safety-relevant structure during model adaptation. Our approach constrains updates in hidden representation space, ensuring that safety-mediating components remain stable while allowing task-specific learning in complementary directions. We evaluate REFUSALGUARD across multiple model families, including LLaMA, Gemma, and Qwen, on adversarial safety benchmarks such as AdvBench, DirectHarm4, and JailbreakBench, as well as downstream utility tasks. Our approach achieves attack success rates comparable to base safety-aligned models while maintaining competitive task performance, significantly outperforming baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01913
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
Asif, Sadia
Amiri, Mohammad Mohammadi
Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computation and Language
Cryptography and Security
Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within the model's activation space, how these representations change during fine-tuning and why alignment degrades remains poorly understood. In this work, we investigate the representation-level mechanisms underlying alignment degradation. Our analysis shows that standard fine-tuning induces systematic drift in safety-relevant representations, distorts their geometric structure, and introduces interference between task optimization and safety features. These effects collectively lead to increased harmful compliance. Motivated by these findings, we introduce REFUSALGUARD, a representation-level fine-tuning framework that preserves safety-relevant structure during model adaptation. Our approach constrains updates in hidden representation space, ensuring that safety-mediating components remain stable while allowing task-specific learning in complementary directions. We evaluate REFUSALGUARD across multiple model families, including LLaMA, Gemma, and Qwen, on adversarial safety benchmarks such as AdvBench, DirectHarm4, and JailbreakBench, as well as downstream utility tasks. Our approach achieves attack success rates comparable to base safety-aligned models while maintaining competitive task performance, significantly outperforming baselines.
title RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
topic Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2605.01913