Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zizzo, Giulio, Cornacchia, Giandomenico, Fraser, Kieran, Hameed, Muhammad Zaid, Rawat, Ambrish, Buesser, Beat, Purcell, Mark, Chen, Pin-Yu, Sattigeri, Prasanna, Varshney, Kush |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024)
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024)
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
di: Rawat, Ambrish, et al.
Pubblicazione: (2024)
di: Rawat, Ambrish, et al.
Pubblicazione: (2024)
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
di: Schoepf, Stefan, et al.
Pubblicazione: (2025)
di: Schoepf, Stefan, et al.
Pubblicazione: (2025)
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
Towards Assurance of LLM Adversarial Robustness using Ontology-Driven Argumentation
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)
Granite Guardian
di: Padhi, Inkit, et al.
Pubblicazione: (2024)
di: Padhi, Inkit, et al.
Pubblicazione: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
di: Hung, Kuo-Han, et al.
Pubblicazione: (2024)
di: Hung, Kuo-Han, et al.
Pubblicazione: (2024)
Domain Adaptation for Time series Transformers using One-step fine-tuning
di: Khanal, Subina, et al.
Pubblicazione: (2024)
di: Khanal, Subina, et al.
Pubblicazione: (2024)
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
di: Huang, Yue, et al.
Pubblicazione: (2025)
di: Huang, Yue, et al.
Pubblicazione: (2025)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
di: Nagireddy, Manish, et al.
Pubblicazione: (2024)
di: Nagireddy, Manish, et al.
Pubblicazione: (2024)
Towards a Practical Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via Randomized Smoothing
di: Gibert, Daniel, et al.
Pubblicazione: (2023)
di: Gibert, Daniel, et al.
Pubblicazione: (2023)
Value Alignment from Unstructured Text
di: Padhi, Inkit, et al.
Pubblicazione: (2024)
di: Padhi, Inkit, et al.
Pubblicazione: (2024)
A Robust Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via (De)Randomized Smoothing
di: Gibert, Daniel, et al.
Pubblicazione: (2024)
di: Gibert, Daniel, et al.
Pubblicazione: (2024)
Differentially Private and Adversarially Robust Machine Learning: An Empirical Evaluation
di: Thakkar, Janvi, et al.
Pubblicazione: (2024)
di: Thakkar, Janvi, et al.
Pubblicazione: (2024)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
di: Thakkar, Janvi, et al.
Pubblicazione: (2023)
di: Thakkar, Janvi, et al.
Pubblicazione: (2023)
Agentic AI Needs a Systems Theory
di: Miehling, Erik, et al.
Pubblicazione: (2025)
di: Miehling, Erik, et al.
Pubblicazione: (2025)
Activated LoRA: Fine-tuned LLMs for Intrinsics
di: Greenewald, Kristjan, et al.
Pubblicazione: (2025)
di: Greenewald, Kristjan, et al.
Pubblicazione: (2025)
Poison Attacks and Adversarial Prompts Against an Informed University Virtual Assistant
di: Fernandez, Ivan A., et al.
Pubblicazione: (2024)
di: Fernandez, Ivan A., et al.
Pubblicazione: (2024)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
di: Young, Richard J.
Pubblicazione: (2025)
di: Young, Richard J.
Pubblicazione: (2025)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
di: Varshney, Kush R.
Pubblicazione: (2023)
di: Varshney, Kush R.
Pubblicazione: (2023)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
di: Varshney, Kush R.
Pubblicazione: (2025)
di: Varshney, Kush R.
Pubblicazione: (2025)
An Algebraic Exposition of the Theory of Dyadic Morality
di: Varshney, Kush R.
Pubblicazione: (2026)
di: Varshney, Kush R.
Pubblicazione: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
di: Islam, Md Asiful, et al.
Pubblicazione: (2026)
di: Islam, Md Asiful, et al.
Pubblicazione: (2026)
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks
di: Mu, Lin, et al.
Pubblicazione: (2025)
di: Mu, Lin, et al.
Pubblicazione: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
di: Chen, Yulin, et al.
Pubblicazione: (2024)
di: Chen, Yulin, et al.
Pubblicazione: (2024)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
di: Wang, Jackson
Pubblicazione: (2026)
di: Wang, Jackson
Pubblicazione: (2026)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
di: Hackett, William, et al.
Pubblicazione: (2025)
di: Hackett, William, et al.
Pubblicazione: (2025)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
Defenses Against Prompt Attacks Learn Surface Heuristics
di: Li, Shawn, et al.
Pubblicazione: (2026)
di: Li, Shawn, et al.
Pubblicazione: (2026)
Prompt Stealing Attacks Against Large Language Models
di: Sha, Zeyang, et al.
Pubblicazione: (2024)
di: Sha, Zeyang, et al.
Pubblicazione: (2024)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
di: Xu, Yue, et al.
Pubblicazione: (2024)
di: Xu, Yue, et al.
Pubblicazione: (2024)
Design Patterns for Securing LLM Agents against Prompt Injections
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2025)
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024) -
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
di: Rawat, Ambrish, et al.
Pubblicazione: (2024) -
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
di: Schoepf, Stefan, et al.
Pubblicazione: (2025) -
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024) -
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
di: Momcilovic, Tomas Bueno, et al.
Pubblicazione: (2024)