Strategic Deflection: Defending LLMs from Logit Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908471324049408 |
|---|---|
| author | Rachidy, Yassine Rbaiti, Jihad Hmamouche, Youssef Sehbaoui, Faissal Seghrouchni, Amal El Fallah |
| author_facet | Rachidy, Yassine Rbaiti, Jihad Hmamouche, Youssef Sehbaoui, Faissal Seghrouchni, Amal El Fallah |
| contents | With the growing adoption of Large Language Models (LLMs) in critical areas, ensuring their security against jailbreaking attacks is paramount. While traditional defenses primarily rely on refusing malicious prompts, recent logit-level attacks have demonstrated the ability to bypass these safeguards by directly manipulating the token-selection process during generation. We introduce Strategic Deflection (SDeflection), a defense that redefines the LLM's response to such advanced attacks. Instead of outright refusal, the model produces an answer that is semantically adjacent to the user's request yet strips away the harmful intent, thereby neutralizing the attacker's harmful intent. Our experiments demonstrate that SDeflection significantly lowers Attack Success Rate (ASR) while maintaining model performance on benign queries. This work presents a critical shift in defensive strategies, moving from simple refusal to strategic content redirection to neutralize advanced threats. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_22160 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Strategic Deflection: Defending LLMs from Logit Manipulation Rachidy, Yassine Rbaiti, Jihad Hmamouche, Youssef Sehbaoui, Faissal Seghrouchni, Amal El Fallah Cryptography and Security Artificial Intelligence Computation and Language With the growing adoption of Large Language Models (LLMs) in critical areas, ensuring their security against jailbreaking attacks is paramount. While traditional defenses primarily rely on refusing malicious prompts, recent logit-level attacks have demonstrated the ability to bypass these safeguards by directly manipulating the token-selection process during generation. We introduce Strategic Deflection (SDeflection), a defense that redefines the LLM's response to such advanced attacks. Instead of outright refusal, the model produces an answer that is semantically adjacent to the user's request yet strips away the harmful intent, thereby neutralizing the attacker's harmful intent. Our experiments demonstrate that SDeflection significantly lowers Attack Success Rate (ASR) while maintaining model performance on benign queries. This work presents a critical shift in defensive strategies, moving from simple refusal to strategic content redirection to neutralize advanced threats. |
| title | Strategic Deflection: Defending LLMs from Logit Manipulation |
| topic | Cryptography and Security Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2507.22160 |