Noise Injection Systemically Degrades Large Language Model Safety Guardrails

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shahani, Prithviraj Singh, Miandoab, Kaveh Eskandari, Scheutz, Matthias
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918158998175744
author Shahani, Prithviraj Singh
Miandoab, Kaveh Eskandari
Scheutz, Matthias
author_facet Shahani, Prithviraj Singh
Miandoab, Kaveh Eskandari
Scheutz, Matthias
contents Safety guardrails in large language models (LLMs) are a critical component in preventing harmful outputs. Yet, their resilience under perturbation remains poorly understood. In this paper, we investigate the robustness of safety fine-tuning in LLMs by systematically injecting Gaussian noise into model activations. We show across multiple open-weight models that (1) Gaussian noise raises harmful-output rates (p < 0.001) by up to 27%, (2) that deeper safety fine-tuning affords no extra protection, and (3) that chain-of-thought reasoning remains largely intact. The findings reveal critical vulnerabilities in current safety alignment techniques and highlight the potential of reasoning-based and reinforcement learning approaches as promising direction for developing more robust AI safety systems. These results have important implications for real-world deployment of LLMs in safety-critical applications as these results imply that widely-deployed safety tuning methods can fail even without adversarial prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Noise Injection Systemically Degrades Large Language Model Safety Guardrails
Shahani, Prithviraj Singh
Miandoab, Kaveh Eskandari
Scheutz, Matthias
Computation and Language
Artificial Intelligence
Machine Learning
Safety guardrails in large language models (LLMs) are a critical component in preventing harmful outputs. Yet, their resilience under perturbation remains poorly understood. In this paper, we investigate the robustness of safety fine-tuning in LLMs by systematically injecting Gaussian noise into model activations. We show across multiple open-weight models that (1) Gaussian noise raises harmful-output rates (p < 0.001) by up to 27%, (2) that deeper safety fine-tuning affords no extra protection, and (3) that chain-of-thought reasoning remains largely intact. The findings reveal critical vulnerabilities in current safety alignment techniques and highlight the potential of reasoning-based and reinforcement learning approaches as promising direction for developing more robust AI safety systems. These results have important implications for real-world deployment of LLMs in safety-critical applications as these results imply that widely-deployed safety tuning methods can fail even without adversarial prompts.
title Noise Injection Systemically Degrades Large Language Model Safety Guardrails
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.13500