Probing the Robustness of Large Language Models Safety to Latent Perturbations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Tianle, Huang, Kexin, Wang, Zongqi, Wang, Yixu, Li, Jie, Yao, Yuanqi, Yao, Yang, Yang, Yujiu, Teng, Yan, Wang, Yingchun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911013752799232
author Gu, Tianle
Huang, Kexin
Wang, Zongqi
Wang, Yixu
Li, Jie
Yao, Yuanqi
Yao, Yang
Yang, Yujiu
Teng, Yan
Wang, Yingchun
author_facet Gu, Tianle
Huang, Kexin
Wang, Zongqi
Wang, Yixu
Li, Jie
Yao, Yuanqi
Yao, Yang
Yang, Yujiu
Teng, Yan
Wang, Yingchun
contents Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that this stems from the shallow nature of existing alignment methods, which focus on surface-level refusal behaviors without sufficiently altering internal representations. Consequently, small shifts in hidden activations can re-trigger harmful behaviors embedded in the latent space. To explore the robustness of safety alignment to latent perturbations, we introduce a probing method that measures the Negative Log-Likelihood of the original response generated by the model. This probe quantifies local sensitivity in the latent space, serving as a diagnostic tool for identifying vulnerable directions. Based on this signal, we construct effective jailbreak trajectories, giving rise to the Activation Steering Attack (ASA). More importantly, these insights offer a principled foundation for improving alignment robustness. To this end, we introduce Layer-wise Adversarial Patch Training~(LAPT), a fine-tuning strategy that inject controlled perturbations into hidden representations during training. Experimental results highlight that LAPT strengthen alignment robustness without compromising general capabilities. Our findings reveal fundamental flaws in current alignment paradigms and call for representation-level training strategies that move beyond surface-level behavior supervision. Codes and results are available at https://github.com/Carol-gutianle/LatentSafety.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing the Robustness of Large Language Models Safety to Latent Perturbations
Gu, Tianle
Huang, Kexin
Wang, Zongqi
Wang, Yixu
Li, Jie
Yao, Yuanqi
Yao, Yang
Yang, Yujiu
Teng, Yan
Wang, Yingchun
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that this stems from the shallow nature of existing alignment methods, which focus on surface-level refusal behaviors without sufficiently altering internal representations. Consequently, small shifts in hidden activations can re-trigger harmful behaviors embedded in the latent space. To explore the robustness of safety alignment to latent perturbations, we introduce a probing method that measures the Negative Log-Likelihood of the original response generated by the model. This probe quantifies local sensitivity in the latent space, serving as a diagnostic tool for identifying vulnerable directions. Based on this signal, we construct effective jailbreak trajectories, giving rise to the Activation Steering Attack (ASA). More importantly, these insights offer a principled foundation for improving alignment robustness. To this end, we introduce Layer-wise Adversarial Patch Training~(LAPT), a fine-tuning strategy that inject controlled perturbations into hidden representations during training. Experimental results highlight that LAPT strengthen alignment robustness without compromising general capabilities. Our findings reveal fundamental flaws in current alignment paradigms and call for representation-level training strategies that move beyond surface-level behavior supervision. Codes and results are available at https://github.com/Carol-gutianle/LatentSafety.
title Probing the Robustness of Large Language Models Safety to Latent Perturbations
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2506.16078