Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Hackett, William, Birch, Lewis, Trawicki, Stefan, Suri, Neeraj, Garraghan, Peter |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
di: Young, Richard J.
Pubblicazione: (2025)
di: Young, Richard J.
Pubblicazione: (2025)
Detecting Prompt Injection Attacks Against Application Using Classifiers
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
di: Pai, Aaditya
Pubblicazione: (2026)
di: Pai, Aaditya
Pubblicazione: (2026)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
di: Shaheer, Safwan, et al.
Pubblicazione: (2025)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
di: Xin, Yuan, et al.
Pubblicazione: (2025)
di: Xin, Yuan, et al.
Pubblicazione: (2025)
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
di: Adiletta, Andrew, et al.
Pubblicazione: (2025)
di: Adiletta, Andrew, et al.
Pubblicazione: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
di: Wei, Xingyuan, et al.
Pubblicazione: (2024)
di: Wei, Xingyuan, et al.
Pubblicazione: (2024)
Defending against Backdoor Attacks via Module Switching
di: Li, Weijun, et al.
Pubblicazione: (2025)
di: Li, Weijun, et al.
Pubblicazione: (2025)
Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design
di: Charles, Richard M., et al.
Pubblicazione: (2025)
di: Charles, Richard M., et al.
Pubblicazione: (2025)
Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models
di: Del Rosario, Ron F.
Pubblicazione: (2025)
di: Del Rosario, Ron F.
Pubblicazione: (2025)
DWFS-Obfuscation: Dynamic Weighted Feature Selection for Robust Malware Familial Classification under Obfuscation
di: Wei, Xingyuan, et al.
Pubblicazione: (2025)
di: Wei, Xingyuan, et al.
Pubblicazione: (2025)
POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
di: Shao, Yangguang, et al.
Pubblicazione: (2025)
di: Shao, Yangguang, et al.
Pubblicazione: (2025)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
di: Wang, Yanshu, et al.
Pubblicazione: (2026)
di: Wang, Yanshu, et al.
Pubblicazione: (2026)
ConfusionPrompt: Practical Private Inference for Online Large Language Models
di: Mai, Peihua, et al.
Pubblicazione: (2023)
di: Mai, Peihua, et al.
Pubblicazione: (2023)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
di: Patlan, Atharv Singh, et al.
Pubblicazione: (2025)
Prompted Contextual Vectors for Spear-Phishing Detection
di: Nahmias, Daniel, et al.
Pubblicazione: (2024)
di: Nahmias, Daniel, et al.
Pubblicazione: (2024)
sudoLLM: On Multi-role Alignment of Language Models
di: Saha, Soumadeep, et al.
Pubblicazione: (2025)
di: Saha, Soumadeep, et al.
Pubblicazione: (2025)
Efficient LLM Safety Evaluation through Multi-Agent Debate
di: Lin, Dachuan, et al.
Pubblicazione: (2025)
di: Lin, Dachuan, et al.
Pubblicazione: (2025)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
di: Wang, Xinhai, et al.
Pubblicazione: (2026)
di: Wang, Xinhai, et al.
Pubblicazione: (2026)
Reducing Information Overload: Because Even Security Experts Need to Blink
di: Kuehn, Philipp, et al.
Pubblicazione: (2022)
di: Kuehn, Philipp, et al.
Pubblicazione: (2022)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
di: Zhang, Peiyan, et al.
Pubblicazione: (2025)
di: Zhang, Peiyan, et al.
Pubblicazione: (2025)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
di: Kurian, Ashley, et al.
Pubblicazione: (2025)
di: Kurian, Ashley, et al.
Pubblicazione: (2025)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
di: Pan, Leyi, et al.
Pubblicazione: (2024)
di: Pan, Leyi, et al.
Pubblicazione: (2024)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
di: Zhang, Tian, et al.
Pubblicazione: (2026)
di: Zhang, Tian, et al.
Pubblicazione: (2026)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
di: Wang, Haochuan Kevin, et al.
Pubblicazione: (2026)
di: Wang, Haochuan Kevin, et al.
Pubblicazione: (2026)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
di: Seddik, Issam, et al.
Pubblicazione: (2025)
di: Seddik, Issam, et al.
Pubblicazione: (2025)
SecEmb: Sparsity-Aware Secure Federated Learning of On-Device Recommender System with Large Embedding
di: Mai, Peihua, et al.
Pubblicazione: (2025)
di: Mai, Peihua, et al.
Pubblicazione: (2025)
Split-and-Denoise: Protect large language model inference with local differential privacy
di: Mai, Peihua, et al.
Pubblicazione: (2023)
di: Mai, Peihua, et al.
Pubblicazione: (2023)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
di: Verma, Apurv, et al.
Pubblicazione: (2024)
di: Verma, Apurv, et al.
Pubblicazione: (2024)
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
di: Kumar, Prasanna
Pubblicazione: (2026)
di: Kumar, Prasanna
Pubblicazione: (2026)
Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts
di: Downey-Webb, Tiarnaigh, et al.
Pubblicazione: (2025)
di: Downey-Webb, Tiarnaigh, et al.
Pubblicazione: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)
di: Othman, Refat
Pubblicazione: (2026)
Prompt Fencing: A Cryptographic Approach to Establishing Security Boundaries in Large Language Model Prompts
di: Peh, Steven
Pubblicazione: (2025)
di: Peh, Steven
Pubblicazione: (2025)
JavelinGuard: Low-Cost Transformer Architectures for LLM Security
di: Datta, Yash, et al.
Pubblicazione: (2025)
di: Datta, Yash, et al.
Pubblicazione: (2025)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
di: Miculicich, Lesly, et al.
Pubblicazione: (2025)
di: Miculicich, Lesly, et al.
Pubblicazione: (2025)
Lightweight LLMs for Network Attack Detection in IoT Networks
di: Sudasinghe, Piyumi Bhagya, et al.
Pubblicazione: (2026)
di: Sudasinghe, Piyumi Bhagya, et al.
Pubblicazione: (2026)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
The Quantum State Continuity Problem and Temporal Enforcement Against Fork Attacks
di: Ünsal, Samet
Pubblicazione: (2025)
di: Ünsal, Samet
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
di: Young, Richard J.
Pubblicazione: (2025) -
Detecting Prompt Injection Attacks Against Application Using Classifiers
di: Shaheer, Safwan, et al.
Pubblicazione: (2025) -
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
di: Pai, Aaditya
Pubblicazione: (2026) -
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
di: Shaheer, Safwan, et al.
Pubblicazione: (2025) -
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
di: Xin, Yuan, et al.
Pubblicazione: (2025)