Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
Fuente:
arXiv
Salvato in:
| Autori principali: | Bertollo, Giacomo, Bodemir, Naz, Burgess, Jonah |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enhancing Guardrails for Safe and Secure Healthcare AI
di: Gangavarapu, Ananya
Pubblicazione: (2024)
di: Gangavarapu, Ananya
Pubblicazione: (2024)
A Comparative Evaluation of AI Agent Security Guardrails
di: Li, Qi, et al.
Pubblicazione: (2026)
di: Li, Qi, et al.
Pubblicazione: (2026)
No Free Lunch with Guardrails
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
Provably Secure Agent Guardrail
di: Wu, Benlong, et al.
Pubblicazione: (2026)
di: Wu, Benlong, et al.
Pubblicazione: (2026)
DNN-Defender: A Victim-Focused In-DRAM Defense Mechanism for Taming Adversarial Weight Attack on DNNs
di: Zhou, Ranyang, et al.
Pubblicazione: (2023)
di: Zhou, Ranyang, et al.
Pubblicazione: (2023)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
di: Wong, Ryan, et al.
Pubblicazione: (2025)
di: Wong, Ryan, et al.
Pubblicazione: (2025)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
di: Jin, Xisen, et al.
Pubblicazione: (2026)
di: Jin, Xisen, et al.
Pubblicazione: (2026)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
di: Wang, Xunguang, et al.
Pubblicazione: (2024)
di: Wang, Xunguang, et al.
Pubblicazione: (2024)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
di: Casper, Stephen, et al.
Pubblicazione: (2024)
di: Casper, Stephen, et al.
Pubblicazione: (2024)
Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7
di: Aydin, Yuksel
Pubblicazione: (2025)
di: Aydin, Yuksel
Pubblicazione: (2025)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
di: Wang, Xunguang, et al.
Pubblicazione: (2025)
di: Wang, Xunguang, et al.
Pubblicazione: (2025)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
di: Li, Nanxi, et al.
Pubblicazione: (2026)
di: Li, Nanxi, et al.
Pubblicazione: (2026)
Current state of LLM Risks and AI Guardrails
di: Ayyamperumal, Suriya Ganesh, et al.
Pubblicazione: (2024)
di: Ayyamperumal, Suriya Ganesh, et al.
Pubblicazione: (2024)
CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks
di: Patil, KrishnaSaiReddy
Pubblicazione: (2026)
di: Patil, KrishnaSaiReddy
Pubblicazione: (2026)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
di: Wu, ChenYu, et al.
Pubblicazione: (2025)
di: Wu, ChenYu, et al.
Pubblicazione: (2025)
AgentWall: A Runtime Safety Layer for Local AI Agents
di: Aravind, Ashwin
Pubblicazione: (2026)
di: Aravind, Ashwin
Pubblicazione: (2026)
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
di: Kang, Daewon, et al.
Pubblicazione: (2025)
di: Kang, Daewon, et al.
Pubblicazione: (2025)
The End of Trust: How Agentic AI Breaks Security Assumptions
di: Zafar, Osama, et al.
Pubblicazione: (2026)
di: Zafar, Osama, et al.
Pubblicazione: (2026)
Defending against Stegomalware in Deep Neural Networks with Permutation Symmetry
di: Torpmann-Hagen, Birk, et al.
Pubblicazione: (2025)
di: Torpmann-Hagen, Birk, et al.
Pubblicazione: (2025)
MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models
di: Cheng, Xueqi, et al.
Pubblicazione: (2025)
di: Cheng, Xueqi, et al.
Pubblicazione: (2025)
Defending Against Beta Poisoning Attacks in Machine Learning Models
di: Gulciftci, Nilufer, et al.
Pubblicazione: (2025)
di: Gulciftci, Nilufer, et al.
Pubblicazione: (2025)
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
di: Tsai, Yu-Che, et al.
Pubblicazione: (2026)
di: Tsai, Yu-Che, et al.
Pubblicazione: (2026)
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
di: Xue, Zhiyu, et al.
Pubblicazione: (2024)
di: Xue, Zhiyu, et al.
Pubblicazione: (2024)
In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
di: Durner, Nils
Pubblicazione: (2025)
di: Durner, Nils
Pubblicazione: (2025)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
Quantifying and Defending against Privacy Threats on Federated Knowledge Graph Embedding
di: Hu, Yuke, et al.
Pubblicazione: (2023)
di: Hu, Yuke, et al.
Pubblicazione: (2023)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
di: Campbell, David, et al.
Pubblicazione: (2026)
di: Campbell, David, et al.
Pubblicazione: (2026)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
di: Zhu, Hongyu, et al.
Pubblicazione: (2024)
di: Zhu, Hongyu, et al.
Pubblicazione: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
OneShield -- the Next Generation of LLM Guardrails
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
di: Jin, Haotian, et al.
Pubblicazione: (2025)
di: Jin, Haotian, et al.
Pubblicazione: (2025)
BitAbuse: A Dataset of Visually Perturbed Texts for Defending Phishing Attacks
di: Lee, Hanyong, et al.
Pubblicazione: (2025)
di: Lee, Hanyong, et al.
Pubblicazione: (2025)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
di: Zhang, Ruyi, et al.
Pubblicazione: (2026)
di: Zhang, Ruyi, et al.
Pubblicazione: (2026)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
What Breaks Embodied AI Security:LLM Vulnerabilities, CPS Flaws,or Something Else?
di: Ma, Boyang, et al.
Pubblicazione: (2026)
di: Ma, Boyang, et al.
Pubblicazione: (2026)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2026)
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2026)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
di: Pang, Yan, et al.
Pubblicazione: (2025)
di: Pang, Yan, et al.
Pubblicazione: (2025)
Uplifted Attackers, Human Defenders: The Cyber Offense-Defense Balance for Trailing-Edge Organizations
di: Murphy, Benjamin, et al.
Pubblicazione: (2025)
di: Murphy, Benjamin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Enhancing Guardrails for Safe and Secure Healthcare AI
di: Gangavarapu, Ananya
Pubblicazione: (2024) -
A Comparative Evaluation of AI Agent Security Guardrails
di: Li, Qi, et al.
Pubblicazione: (2026) -
No Free Lunch with Guardrails
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025) -
Provably Secure Agent Guardrail
di: Wu, Benlong, et al.
Pubblicazione: (2026) -
DNN-Defender: A Victim-Focused In-DRAM Defense Mechanism for Taming Adversarial Weight Attack on DNNs
di: Zhou, Ranyang, et al.
Pubblicazione: (2023)