Strategic Deflection: Defending LLMs from Logit Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rachidy, Yassine, Rbaiti, Jihad, Hmamouche, Youssef, Sehbaoui, Faissal, Seghrouchni, Amal El Fallah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908471324049408
author Rachidy, Yassine
Rbaiti, Jihad
Hmamouche, Youssef
Sehbaoui, Faissal
Seghrouchni, Amal El Fallah
author_facet Rachidy, Yassine
Rbaiti, Jihad
Hmamouche, Youssef
Sehbaoui, Faissal
Seghrouchni, Amal El Fallah
contents With the growing adoption of Large Language Models (LLMs) in critical areas, ensuring their security against jailbreaking attacks is paramount. While traditional defenses primarily rely on refusing malicious prompts, recent logit-level attacks have demonstrated the ability to bypass these safeguards by directly manipulating the token-selection process during generation. We introduce Strategic Deflection (SDeflection), a defense that redefines the LLM's response to such advanced attacks. Instead of outright refusal, the model produces an answer that is semantically adjacent to the user's request yet strips away the harmful intent, thereby neutralizing the attacker's harmful intent. Our experiments demonstrate that SDeflection significantly lowers Attack Success Rate (ASR) while maintaining model performance on benign queries. This work presents a critical shift in defensive strategies, moving from simple refusal to strategic content redirection to neutralize advanced threats.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Strategic Deflection: Defending LLMs from Logit Manipulation
Rachidy, Yassine
Rbaiti, Jihad
Hmamouche, Youssef
Sehbaoui, Faissal
Seghrouchni, Amal El Fallah
Cryptography and Security
Artificial Intelligence
Computation and Language
With the growing adoption of Large Language Models (LLMs) in critical areas, ensuring their security against jailbreaking attacks is paramount. While traditional defenses primarily rely on refusing malicious prompts, recent logit-level attacks have demonstrated the ability to bypass these safeguards by directly manipulating the token-selection process during generation. We introduce Strategic Deflection (SDeflection), a defense that redefines the LLM's response to such advanced attacks. Instead of outright refusal, the model produces an answer that is semantically adjacent to the user's request yet strips away the harmful intent, thereby neutralizing the attacker's harmful intent. Our experiments demonstrate that SDeflection significantly lowers Attack Success Rate (ASR) while maintaining model performance on benign queries. This work presents a critical shift in defensive strategies, moving from simple refusal to strategic content redirection to neutralize advanced threats.
title Strategic Deflection: Defending LLMs from Logit Manipulation
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.22160