Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Armstrong, Stuart, Franklin, Matija, Stevens, Connor, Gorman, Rebecca |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Combating Phone Scams with LLM-based Detection: Where Do We Stand?
par: Shen, Zitong, et autres
Publié: (2024)
par: Shen, Zitong, et autres
Publié: (2024)
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
par: Ntais, Pavlos
Publié: (2025)
par: Ntais, Pavlos
Publié: (2025)
PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices
par: Chattopadhyay, Nandish, et autres
Publié: (2026)
par: Chattopadhyay, Nandish, et autres
Publié: (2026)
Browser Extension for Fake URL Detection
par: Malik, Latesh G., et autres
Publié: (2024)
par: Malik, Latesh G., et autres
Publié: (2024)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
par: Wang, Yanshu, et autres
Publié: (2026)
par: Wang, Yanshu, et autres
Publié: (2026)
Data Defenses Against Large Language Models
par: Agnew, William, et autres
Publié: (2024)
par: Agnew, William, et autres
Publié: (2024)
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
par: Tshimula, Jean Marie, et autres
Publié: (2024)
par: Tshimula, Jean Marie, et autres
Publié: (2024)
h4rm3l: A language for Composable Jailbreak Attack Synthesis
par: Doumbouya, Moussa Koulako Bala, et autres
Publié: (2024)
par: Doumbouya, Moussa Koulako Bala, et autres
Publié: (2024)
David and Goliath: An Empirical Evaluation of Attacks and Defenses for QNNs at the Deep Edge
par: Costa, Miguel, et autres
Publié: (2024)
par: Costa, Miguel, et autres
Publié: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
par: Li, Yucheng, et autres
Publié: (2025)
par: Li, Yucheng, et autres
Publié: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
par: Li, Nathaniel, et autres
Publié: (2024)
par: Li, Nathaniel, et autres
Publié: (2024)
Generating Image Adversarial Examples by Embedding Digital Watermarks
par: Xiang, Yuexin, et autres
Publié: (2020)
par: Xiang, Yuexin, et autres
Publié: (2020)
Defense Against Indirect Prompt Injection via Tool Result Parsing
par: Yu, Qiang, et autres
Publié: (2026)
par: Yu, Qiang, et autres
Publié: (2026)
AI Agents May Always Fall for Prompt Injections
par: Abdelnabi, Sahar, et autres
Publié: (2026)
par: Abdelnabi, Sahar, et autres
Publié: (2026)
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
par: Sekar, Anirudh, et autres
Publié: (2026)
par: Sekar, Anirudh, et autres
Publié: (2026)
Sparse vs Contiguous Adversarial Pixel Perturbations in Multimodal Models: An Empirical Analysis
par: Botocan, Cristian-Alexandru, et autres
Publié: (2024)
par: Botocan, Cristian-Alexandru, et autres
Publié: (2024)
Deep Learning-Based Intrusion Detection for Automotive Ethernet: Evaluating & Optimizing Fast Inference Techniques for Deployment on Low-Cost Platform
par: Carmo, Pedro R. X., et autres
Publié: (2025)
par: Carmo, Pedro R. X., et autres
Publié: (2025)
Scalable and Ethical Insider Threat Detection through Data Synthesis and Analysis by LLMs
par: Gelman, Haywood, et autres
Publié: (2025)
par: Gelman, Haywood, et autres
Publié: (2025)
An Ethically Grounded LLM-Based Approach to Insider Threat Synthesis and Detection
par: Gelman, Haywood, et autres
Publié: (2025)
par: Gelman, Haywood, et autres
Publié: (2025)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
par: Liu, Houjun, et autres
Publié: (2026)
par: Liu, Houjun, et autres
Publié: (2026)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
par: Correia, Pedro H. Barcha, et autres
Publié: (2026)
par: Correia, Pedro H. Barcha, et autres
Publié: (2026)
Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts
par: Downey-Webb, Tiarnaigh, et autres
Publié: (2025)
par: Downey-Webb, Tiarnaigh, et autres
Publié: (2025)
AI-Driven Security Alert Screening and Alert Fatigue Mitigation in Security Operations Centers: A Survey
par: Ndichu, Samuel, et autres
Publié: (2026)
par: Ndichu, Samuel, et autres
Publié: (2026)
Organizational Adaptation to Generative AI in Cybersecurity
par: Nott, Christopher
Publié: (2025)
par: Nott, Christopher
Publié: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
par: Pathade, Chetan
Publié: (2025)
par: Pathade, Chetan
Publié: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
par: Zhang, Chiyu, et autres
Publié: (2025)
par: Zhang, Chiyu, et autres
Publié: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
par: Chen, Yunhao, et autres
Publié: (2025)
par: Chen, Yunhao, et autres
Publié: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
par: An, Bang, et autres
Publié: (2024)
par: An, Bang, et autres
Publié: (2024)
Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity
par: Nguyen, Tam n.
Publié: (2024)
par: Nguyen, Tam n.
Publié: (2024)
TSFool: Crafting Highly-Imperceptible Adversarial Time Series through Multi-Objective Attack
par: Wang, Yanyun, et autres
Publié: (2022)
par: Wang, Yanyun, et autres
Publié: (2022)
Evaluating Large Language Models for Causal Modeling
par: Razouk, Houssam, et autres
Publié: (2024)
par: Razouk, Houssam, et autres
Publié: (2024)
Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks
par: Parmar, Manoj
Publié: (2026)
par: Parmar, Manoj
Publié: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
par: Murphy, Brendan, et autres
Publié: (2025)
par: Murphy, Brendan, et autres
Publié: (2025)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
par: Wang, Jiawen, et autres
Publié: (2025)
par: Wang, Jiawen, et autres
Publié: (2025)
LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
par: Yang, Zhuoran, et autres
Publié: (2025)
par: Yang, Zhuoran, et autres
Publié: (2025)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
par: Mo, Yichuan, et autres
Publié: (2024)
par: Mo, Yichuan, et autres
Publié: (2024)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
par: Li, Linbao, et autres
Publié: (2025)
par: Li, Linbao, et autres
Publié: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
par: Yu, Zhiyuan, et autres
Publié: (2024)
par: Yu, Zhiyuan, et autres
Publié: (2024)
A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty
par: Lin, Zehao, et autres
Publié: (2026)
par: Lin, Zehao, et autres
Publié: (2026)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
par: Chen, Bocheng, et autres
Publié: (2024)
par: Chen, Bocheng, et autres
Publié: (2024)
Documents similaires
-
Combating Phone Scams with LLM-based Detection: Where Do We Stand?
par: Shen, Zitong, et autres
Publié: (2024) -
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
par: Ntais, Pavlos
Publié: (2025) -
PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices
par: Chattopadhyay, Nandish, et autres
Publié: (2026) -
Browser Extension for Fake URL Detection
par: Malik, Latesh G., et autres
Publié: (2024) -
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
par: Wang, Yanshu, et autres
Publié: (2026)