Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Tshimula, Jean Marie, Ndona, Xavier, Nkashama, D'Jeff K., Tardif, Pierre-Martin, Kabanza, Froduald, Frappier, Marc, Wang, Shengrui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Learning for Network Anomaly Detection under Data Contamination: Evaluating Robustness and Mitigating Performance Degradation
by: Nkashama, D'Jeff K., et al.
Published: (2024)
by: Nkashama, D'Jeff K., et al.
Published: (2024)
Impact of Inaccurate Contamination Ratio on Robust Unsupervised Anomaly Detection
by: Masakuna, Jordan F., et al.
Published: (2024)
by: Masakuna, Jordan F., et al.
Published: (2024)
Psychological Profiling in Cybersecurity: A Look at LLMs and Psycholinguistic Features
by: Tshimula, Jean Marie, et al.
Published: (2024)
by: Tshimula, Jean Marie, et al.
Published: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
Towards in-situ Psychological Profiling of Cybercriminals Using Dynamically Generated Deception Environments
by: Quibell, Jacob
Published: (2024)
by: Quibell, Jacob
Published: (2024)
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
by: Yang, Guangyu, et al.
Published: (2025)
by: Yang, Guangyu, et al.
Published: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Quantitative Resilience Modeling for Autonomous Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
Infrastructure Patterns in Toll Scam Domains: A Comprehensive Analysis of Cybercriminal Registration and Hosting Strategies
by: Munny, Morium Akter, et al.
Published: (2025)
by: Munny, Morium Akter, et al.
Published: (2025)
MalTool: Malicious Tool Attacks on LLM Agents
by: Hu, Yuepeng, et al.
Published: (2026)
by: Hu, Yuepeng, et al.
Published: (2026)
ASTD Patterns for Integrated Continuous Anomaly Detection In Data Logs
by: Jabri, Chaymae El, et al.
Published: (2024)
by: Jabri, Chaymae El, et al.
Published: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
Cross-Domain AI for Early Attack Detection and Defense Against Malicious Flows in O-RAN
by: Xavier, Bruno Missi, et al.
Published: (2024)
by: Xavier, Bruno Missi, et al.
Published: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
by: Li, Yucheng, et al.
Published: (2025)
by: Li, Yucheng, et al.
Published: (2025)
The Path To Autonomous Cyber Defense
by: Oesch, Sean, et al.
Published: (2024)
by: Oesch, Sean, et al.
Published: (2024)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
by: Zhong, Xingwei, et al.
Published: (2025)
by: Zhong, Xingwei, et al.
Published: (2025)
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
by: Armstrong, Stuart, et al.
Published: (2025)
by: Armstrong, Stuart, et al.
Published: (2025)
Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense
by: Prinos, Kerri, et al.
Published: (2026)
by: Prinos, Kerri, et al.
Published: (2026)
DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on Telegram
by: Roy, Sayak Saha, et al.
Published: (2024)
by: Roy, Sayak Saha, et al.
Published: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
by: Derya, Kemal, et al.
Published: (2026)
by: Derya, Kemal, et al.
Published: (2026)
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
by: Du, Mengyao, et al.
Published: (2026)
by: Du, Mengyao, et al.
Published: (2026)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
by: Gu, Haoran, et al.
Published: (2026)
by: Gu, Haoran, et al.
Published: (2026)
Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
by: Mishra, Rina, et al.
Published: (2025)
by: Mishra, Rina, et al.
Published: (2025)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Firewall Regulatory Networks for Autonomous Cyber Defense
by: Duan, Qi, et al.
Published: (2025)
by: Duan, Qi, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
PoolFlip: A Multi-Agent Reinforcement Learning Security Environment for Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
by: Brundage, Miles, et al.
Published: (2018)
by: Brundage, Miles, et al.
Published: (2018)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
by: Kim, Taeyoun, et al.
Published: (2024)
by: Kim, Taeyoun, et al.
Published: (2024)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
by: Yang, Xianglin, et al.
Published: (2025)
by: Yang, Xianglin, et al.
Published: (2025)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Guardians of the Agentic System: Preventing Many Shots Jailbreak with Agentic System
by: Barua, Saikat, et al.
Published: (2025)
by: Barua, Saikat, et al.
Published: (2025)
LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
by: Yang, Zhuoran, et al.
Published: (2025)
by: Yang, Zhuoran, et al.
Published: (2025)
Similar Items
-
Deep Learning for Network Anomaly Detection under Data Contamination: Evaluating Robustness and Mitigating Performance Degradation
by: Nkashama, D'Jeff K., et al.
Published: (2024) -
Impact of Inaccurate Contamination Ratio on Robust Unsupervised Anomaly Detection
by: Masakuna, Jordan F., et al.
Published: (2024) -
Psychological Profiling in Cybersecurity: A Look at LLMs and Psycholinguistic Features
by: Tshimula, Jean Marie, et al.
Published: (2024) -
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024) -
Towards in-situ Psychological Profiling of Cybercriminals Using Dynamically Generated Deception Environments
by: Quibell, Jacob
Published: (2024)