Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
Fuente:
arXiv
Saved in:
| Main Authors: | Rawat, Ambrish, Schoepf, Stefan, Zizzo, Giulio, Cornacchia, Giandomenico, Hameed, Muhammad Zaid, Fraser, Kieran, Miehling, Erik, Buesser, Beat, Daly, Elizabeth M., Purcell, Mark, Sattigeri, Prasanna, Chen, Pin-Yu, Varshney, Kush R. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
by: Schoepf, Stefan, et al.
Published: (2025)
by: Schoepf, Stefan, et al.
Published: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
by: Cornacchia, Giandomenico, et al.
Published: (2024)
by: Cornacchia, Giandomenico, et al.
Published: (2024)
Granite Guardian
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
Towards Assurance of LLM Adversarial Robustness using Ontology-Driven Argumentation
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Agentic AI Needs a Systems Theory
by: Miehling, Erik, et al.
Published: (2025)
by: Miehling, Erik, et al.
Published: (2025)
Domain Adaptation for Time series Transformers using One-step fine-tuning
by: Khanal, Subina, et al.
Published: (2024)
by: Khanal, Subina, et al.
Published: (2024)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Value Alignment from Unstructured Text
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
Activated LoRA: Fine-tuned LLMs for Intrinsics
by: Greenewald, Kristjan, et al.
Published: (2025)
by: Greenewald, Kristjan, et al.
Published: (2025)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
by: Varshney, Kush R.
Published: (2023)
by: Varshney, Kush R.
Published: (2023)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
by: Varshney, Kush R.
Published: (2025)
by: Varshney, Kush R.
Published: (2025)
An Algebraic Exposition of the Theory of Dyadic Morality
by: Varshney, Kush R.
Published: (2026)
by: Varshney, Kush R.
Published: (2026)
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
by: Bagehorn, Frank, et al.
Published: (2025)
by: Bagehorn, Frank, et al.
Published: (2025)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Gamma, Gaussian and Poisson approximations for random sums using size-biased and generalized zero-biased couplings
by: Daly, Fraser
Published: (2020)
by: Daly, Fraser
Published: (2020)
Approximations for the number of maxima and near-maxima in independent data
by: Daly, Fraser
Published: (2025)
by: Daly, Fraser
Published: (2025)
On Stein's method for stochastically monotone single-birth chains
by: Daly, Fraser
Published: (2022)
by: Daly, Fraser
Published: (2022)
Biasing with an independent increment: Gaussian approximations and proximity of Poisson mixtures
by: Daly, Fraser
Published: (2025)
by: Daly, Fraser
Published: (2025)
Poisson and Gaussian approximations of the power divergence family of statistics
by: Daly, Fraser
Published: (2023)
by: Daly, Fraser
Published: (2023)
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
Contextual Moral Value Alignment Through Context-Based Aggregation
by: Dognin, Pierre, et al.
Published: (2024)
by: Dognin, Pierre, et al.
Published: (2024)
Differentially Private and Adversarially Robust Machine Learning: An Empirical Evaluation
by: Thakkar, Janvi, et al.
Published: (2024)
by: Thakkar, Janvi, et al.
Published: (2024)
Towards a Practical Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via Randomized Smoothing
by: Gibert, Daniel, et al.
Published: (2023)
by: Gibert, Daniel, et al.
Published: (2023)
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
by: Belkhiter, Yannis, et al.
Published: (2024)
by: Belkhiter, Yannis, et al.
Published: (2024)
Blue Teaming Function-Calling Agents
by: Dolcetti, Greta, et al.
Published: (2026)
by: Dolcetti, Greta, et al.
Published: (2026)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
by: Thakkar, Janvi, et al.
Published: (2023)
by: Thakkar, Janvi, et al.
Published: (2023)
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
by: Bouneffouf, Djallel, et al.
Published: (2025)
by: Bouneffouf, Djallel, et al.
Published: (2025)
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
by: Nagireddy, Manish, et al.
Published: (2024)
by: Nagireddy, Manish, et al.
Published: (2024)
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025)
by: Cintas, Celia, et al.
Published: (2025)
Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond
by: Ceccon, Marina, et al.
Published: (2025)
by: Ceccon, Marina, et al.
Published: (2025)
On geometric-type approximations with applications
by: Daly, Fraser, et al.
Published: (2023)
by: Daly, Fraser, et al.
Published: (2023)
Rates of convergence for extremal spacings in Kakutani's random interval-splitting process
by: Daly, Fraser, et al.
Published: (2025)
by: Daly, Fraser, et al.
Published: (2025)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Verifiability and Privacy in Federated Learning through Context-Hiding Multi-Key Homomorphic Authenticators
by: Bottoni, Simone, et al.
Published: (2025)
by: Bottoni, Simone, et al.
Published: (2025)
A Robust Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via (De)Randomized Smoothing
by: Gibert, Daniel, et al.
Published: (2024)
by: Gibert, Daniel, et al.
Published: (2024)
Site-specific ILC Detector Installation Plan
by: Buesser, Karsten, et al.
Published: (2026)
by: Buesser, Karsten, et al.
Published: (2026)
Similar Items
-
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025) -
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
by: Schoepf, Stefan, et al.
Published: (2025) -
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
by: Cornacchia, Giandomenico, et al.
Published: (2024) -
Granite Guardian
by: Padhi, Inkit, et al.
Published: (2024) -
Towards Assurance of LLM Adversarial Robustness using Ontology-Driven Argumentation
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)