You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense
Fuente:
arXiv
Saved in:
| Main Authors: | Mai, Wuyuao, Hong, Geng, Chen, Pei, Pan, Xudong, Liu, Baojun, Zhang, Yuan, Duan, Haixin, Yang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
by: Mai, Wuyuao, et al.
Published: (2025)
by: Mai, Wuyuao, et al.
Published: (2025)
Have Your Cake and Eat It Too: Toward Efficient and Accurate Split Federated Learning
by: Yan, Dengke, et al.
Published: (2023)
by: Yan, Dengke, et al.
Published: (2023)
Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration
by: Wu, Mengying, et al.
Published: (2024)
by: Wu, Mengying, et al.
Published: (2024)
You Can't Trust Your Tag Neither: Privacy Leaks and Potential Legal Violations within the Google Tag Manager
by: Mertens, Gilles, et al.
Published: (2023)
by: Mertens, Gilles, et al.
Published: (2023)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025)
by: Cao, Bochuan, et al.
Published: (2025)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
You Can Run But You Can't Hide: Runtime Protection Against Malicious Package Updates For Node.js
by: Ohm, Marc, et al.
Published: (2023)
by: Ohm, Marc, et al.
Published: (2023)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search
by: Shen, Yulin, et al.
Published: (2026)
by: Shen, Yulin, et al.
Published: (2026)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
by: Zheng, Xiaosen, et al.
Published: (2024)
by: Zheng, Xiaosen, et al.
Published: (2024)
PRISON: Unmasking the Criminal Potential of Large Language Models
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
by: Zhang, Yunyi, et al.
Published: (2025)
by: Zhang, Yunyi, et al.
Published: (2025)
If You Can't Beat Them, Eat Them – Stinging Nettle
by: NA
Published: (1989)
by: NA
Published: (1989)
Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent
by: Wu, Xinyi, et al.
Published: (2026)
by: Wu, Xinyi, et al.
Published: (2026)
Feedback-Guided Extraction of Knowledge Base from Retrieval-Augmented LLM Applications
by: Jiang, Changyue, et al.
Published: (2024)
by: Jiang, Changyue, et al.
Published: (2024)
MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent
by: Wu, Mengying, et al.
Published: (2025)
by: Wu, Mengying, et al.
Published: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025)
by: Li, Xiaohu, et al.
Published: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025)
by: Zhong, Xingwei, et al.
Published: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
by: Qiu, Huming, et al.
Published: (2023)
by: Qiu, Huming, et al.
Published: (2023)
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything
by: Zou, Xiaotian, et al.
Published: (2024)
by: Zou, Xiaotian, et al.
Published: (2024)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
by: Mao, Yanxu, et al.
Published: (2025)
by: Mao, Yanxu, et al.
Published: (2025)
Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use
by: Zhang, Wuyang, et al.
Published: (2026)
by: Zhang, Wuyang, et al.
Published: (2026)
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
by: Derya, Kemal, et al.
Published: (2026)
by: Derya, Kemal, et al.
Published: (2026)
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
by: Nikolić, Kristina, et al.
Published: (2025)
by: Nikolić, Kristina, et al.
Published: (2025)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation
by: Chen, Pei, et al.
Published: (2026)
by: Chen, Pei, et al.
Published: (2026)
Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
by: Mishra, Rina, et al.
Published: (2025)
by: Mishra, Rina, et al.
Published: (2025)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Similar Items
-
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
by: Mai, Wuyuao, et al.
Published: (2025) -
Have Your Cake and Eat It Too: Toward Efficient and Accurate Split Federated Learning
by: Yan, Dengke, et al.
Published: (2023) -
Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration
by: Wu, Mengying, et al.
Published: (2024) -
You Can't Trust Your Tag Neither: Privacy Leaks and Potential Legal Violations within the Google Tag Manager
by: Mertens, Gilles, et al.
Published: (2023) -
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025)