Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
Fuente:
arXiv
Saved in:
| Main Authors: | Wong, Ryan, Ng, Hosea David Yu Fei, Sharma, Dhananjai, Ng, Glenn Jun Jie, Srinivasan, Kavishvaran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025)
by: Zhang, Shenyi, et al.
Published: (2025)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025)
by: Yin, Ziyi, et al.
Published: (2025)
AdaPhish: AI-Powered Adaptive Defense and Education Resource Against Deceptive Emails
by: Meguro, Rei, et al.
Published: (2025)
by: Meguro, Rei, et al.
Published: (2025)
Defending Against Prompt Injection with DataFilter
by: Wang, Yizhu, et al.
Published: (2025)
by: Wang, Yizhu, et al.
Published: (2025)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
by: Xu, Zhenhua, et al.
Published: (2026)
by: Xu, Zhenhua, et al.
Published: (2026)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
by: Wu, Daoyuan, et al.
Published: (2024)
by: Wu, Daoyuan, et al.
Published: (2024)
Defending Against Intelligent Attackers at Large Scales
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
StruQ: Defending Against Prompt Injection with Structured Queries
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
Defending Against Prompt Injection With a Few DefensiveTokens
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Defending against Jailbreak through Early Exit Generation of Large Language Models
by: Zhao, Chongwen, et al.
Published: (2024)
by: Zhao, Chongwen, et al.
Published: (2024)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
Defending Large Language Models Against Attacks With Residual Stream Activation Analysis
by: Kawasaki, Amelia, et al.
Published: (2024)
by: Kawasaki, Amelia, et al.
Published: (2024)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024)
by: Jiang, Tanqiu, et al.
Published: (2024)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
by: Li, Yanzeng, et al.
Published: (2025)
by: Li, Yanzeng, et al.
Published: (2025)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
SPML: A DSL for Defending Language Models Against Prompt Attacks
by: Sharma, Reshabh K, et al.
Published: (2024)
by: Sharma, Reshabh K, et al.
Published: (2024)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
TextCrafter: Optimization-Calibrated Noise for Defending Against Text Embedding Inversion
by: Tang, Duoxun, et al.
Published: (2025)
by: Tang, Duoxun, et al.
Published: (2025)
Defending Against Neural Network Model Inversion Attacks via Data Poisoning
by: Zhou, Shuai, et al.
Published: (2024)
by: Zhou, Shuai, et al.
Published: (2024)
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
JULI: Jailbreak Large Language Models by Self-Introspection
by: Wang, Jesson, et al.
Published: (2025)
by: Wang, Jesson, et al.
Published: (2025)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
by: Li, Songze, et al.
Published: (2026)
by: Li, Songze, et al.
Published: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
by: Yadav, Advait, et al.
Published: (2025)
by: Yadav, Advait, et al.
Published: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
by: Li, Jie, et al.
Published: (2024)
by: Li, Jie, et al.
Published: (2024)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)
by: Zhong, Yinan, et al.
Published: (2025)
Secure Distributed Learning for CAVs: Defending Against Gradient Leakage with Leveled Homomorphic Encryption
by: Najjar, Muhammad Ali, et al.
Published: (2025)
by: Najjar, Muhammad Ali, et al.
Published: (2025)
Defending Against Attack on the Cloned: In-Band Active Man-in-the-Middle Detection for the Signal Protocol
by: Teng, Wil Liam, et al.
Published: (2024)
by: Teng, Wil Liam, et al.
Published: (2024)
Similar Items
-
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024) -
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025) -
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025) -
AdaPhish: AI-Powered Adaptive Defense and Education Resource Against Deceptive Emails
by: Meguro, Rei, et al.
Published: (2025)