Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wong, Ryan, Ng, Hosea David Yu Fei, Sharma, Dhananjai, Ng, Glenn Jun Jie, Srinivasan, Kavishvaran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915634989760512
author Wong, Ryan
Ng, Hosea David Yu Fei
Sharma, Dhananjai
Ng, Glenn Jun Jie
Srinivasan, Kavishvaran
author_facet Wong, Ryan
Ng, Hosea David Yu Fei
Sharma, Dhananjai
Ng, Glenn Jun Jie
Srinivasan, Kavishvaran
contents Large Language Models (LLMs) remain susceptible to jailbreak exploits that bypass safety filters and induce harmful or unethical behavior. This work presents a systematic taxonomy of existing jailbreak defenses across prompt-level, model-level, and training-time interventions, followed by three proposed defense strategies. First, a Prompt-Level Defense Framework detects and neutralizes adversarial inputs through sanitization, paraphrasing, and adaptive system guarding. Second, a Logit-Based Steering Defense reinforces refusal behavior through inference-time vector steering in safety-sensitive layers. Third, a Domain-Specific Agent Defense employs the MetaGPT framework to enforce structured, role-based collaboration and domain adherence. Experiments on benchmark datasets show substantial reductions in attack success rate, achieving full mitigation under the agent-based defense. Overall, this study highlights how jailbreaks pose a significant security threat to LLMs and identifies key intervention points for prevention, while noting that defense strategies often involve trade-offs between safety, performance, and scalability. Code is available at: https://github.com/Kuro0911/CS5446-Project
format Preprint
id arxiv_https___arxiv_org_abs_2511_18933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
Wong, Ryan
Ng, Hosea David Yu Fei
Sharma, Dhananjai
Ng, Glenn Jun Jie
Srinivasan, Kavishvaran
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) remain susceptible to jailbreak exploits that bypass safety filters and induce harmful or unethical behavior. This work presents a systematic taxonomy of existing jailbreak defenses across prompt-level, model-level, and training-time interventions, followed by three proposed defense strategies. First, a Prompt-Level Defense Framework detects and neutralizes adversarial inputs through sanitization, paraphrasing, and adaptive system guarding. Second, a Logit-Based Steering Defense reinforces refusal behavior through inference-time vector steering in safety-sensitive layers. Third, a Domain-Specific Agent Defense employs the MetaGPT framework to enforce structured, role-based collaboration and domain adherence. Experiments on benchmark datasets show substantial reductions in attack success rate, achieving full mitigation under the agent-based defense. Overall, this study highlights how jailbreaks pose a significant security threat to LLMs and identifies key intervention points for prevention, while noting that defense strategies often involve trade-offs between safety, performance, and scalability. Code is available at: https://github.com/Kuro0911/CS5446-Project
title Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2511.18933