Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhibo, Li, Yuxi, Wang, Kailong, Yuan, Shuai, Shi, Ling, Wang, Haoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909683867975680
author Zhang, Zhibo
Li, Yuxi
Wang, Kailong
Yuan, Shuai
Shi, Ling
Wang, Haoyu
author_facet Zhang, Zhibo
Li, Yuxi
Wang, Kailong
Yuan, Shuai
Shi, Ling
Wang, Haoyu
contents Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant security risks, particularly through embedding space poisoning, which is a subtle attack vector where adversaries manipulate the internal semantic representations of input data to bypass safety alignment mechanisms. While previous research has investigated universal perturbation methods, the dynamics of LLM safety alignment at the embedding level remain insufficiently understood. Consequently, more targeted and accurate adversarial perturbation techniques, which pose significant threats, have not been adequately studied. In this work, we propose ETTA (Embedding Transformation Toxicity Attenuation), a novel framework that identifies and attenuates toxicity-sensitive dimensions in embedding space via linear transformations. ETTA bypasses model refusal behaviors while preserving linguistic coherence, without requiring model fine-tuning or access to training data. Evaluated on five representative open-source LLMs using the AdvBench benchmark, ETTA achieves a high average attack success rate of 88.61%, outperforming the best baseline by 11.34%, and generalizes to safety-enhanced models (e.g., 77.39% ASR on instruction-tuned defenses). These results highlight a critical vulnerability in current alignment strategies and underscore the need for embedding-aware defenses.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08020
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
Zhang, Zhibo
Li, Yuxi
Wang, Kailong
Yuan, Shuai
Shi, Ling
Wang, Haoyu
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant security risks, particularly through embedding space poisoning, which is a subtle attack vector where adversaries manipulate the internal semantic representations of input data to bypass safety alignment mechanisms. While previous research has investigated universal perturbation methods, the dynamics of LLM safety alignment at the embedding level remain insufficiently understood. Consequently, more targeted and accurate adversarial perturbation techniques, which pose significant threats, have not been adequately studied. In this work, we propose ETTA (Embedding Transformation Toxicity Attenuation), a novel framework that identifies and attenuates toxicity-sensitive dimensions in embedding space via linear transformations. ETTA bypasses model refusal behaviors while preserving linguistic coherence, without requiring model fine-tuning or access to training data. Evaluated on five representative open-source LLMs using the AdvBench benchmark, ETTA achieves a high average attack success rate of 88.61%, outperforming the best baseline by 11.34%, and generalizes to safety-enhanced models (e.g., 77.39% ASR on instruction-tuned defenses). These results highlight a critical vulnerability in current alignment strategies and underscore the need for embedding-aware defenses.
title Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.08020