Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yinzhi, Wang, Ming, Feng, Shi, Yang, Xiaocui, Wang, Daling, Zhang, Yifei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911410489917440
author Zhao, Yinzhi
Wang, Ming
Feng, Shi
Yang, Xiaocui
Wang, Daling
Zhang, Yifei
author_facet Zhao, Yinzhi
Wang, Ming
Feng, Shi
Yang, Xiaocui
Wang, Daling
Zhang, Yifei
contents Large language models (LLMs) have achieved impressive performance across natural language tasks and are increasingly deployed in real-world applications. Despite extensive safety alignment efforts, recent studies show that such alignment is often shallow and remains vulnerable to jailbreak attacks. Existing defense mechanisms, including decoding-based constraints and post-hoc content detectors, struggle against sophisticated jailbreaks, often intervening robust detection or excessively degrading model utility. In this work, we examine the decoding process of LLMs and make a key observation: even when successfully jailbroken, models internally exhibit latent safety-related signals during generation. However, these signals are overridden by the model's drive for fluent continuation, preventing timely self-correction or refusal. Building on this observation, we propose a simple yet effective approach that explicitly surfaces and leverages these latent safety signals for early detection of unsafe content during decoding. Experiments across diverse jailbreak attacks demonstrate that our approach significantly enhances safety, while maintaining low over-refusal rates on benign inputs and preserving response quality. Our results suggest that activating intrinsic safety-awareness during decoding offers a promising and complementary direction for defending against jailbreak attacks. Code is available at: https://github.com/zyz13590/SafeProbing.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
Zhao, Yinzhi
Wang, Ming
Feng, Shi
Yang, Xiaocui
Wang, Daling
Zhang, Yifei
Artificial Intelligence
Computation and Language
Large language models (LLMs) have achieved impressive performance across natural language tasks and are increasingly deployed in real-world applications. Despite extensive safety alignment efforts, recent studies show that such alignment is often shallow and remains vulnerable to jailbreak attacks. Existing defense mechanisms, including decoding-based constraints and post-hoc content detectors, struggle against sophisticated jailbreaks, often intervening robust detection or excessively degrading model utility. In this work, we examine the decoding process of LLMs and make a key observation: even when successfully jailbroken, models internally exhibit latent safety-related signals during generation. However, these signals are overridden by the model's drive for fluent continuation, preventing timely self-correction or refusal. Building on this observation, we propose a simple yet effective approach that explicitly surfaces and leverages these latent safety signals for early detection of unsafe content during decoding. Experiments across diverse jailbreak attacks demonstrate that our approach significantly enhances safety, while maintaining low over-refusal rates on benign inputs and preserving response quality. Our results suggest that activating intrinsic safety-awareness during decoding offers a promising and complementary direction for defending against jailbreak attacks. Code is available at: https://github.com/zyz13590/SafeProbing.
title Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.10543