Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Barrett, Anthony M., Jackson, Krystal, Murphy, Evan R., Madkour, Nada, Newman, Jessica |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
Toward Risk Thresholds for AI-Enabled Cyber Threats: Enhancing Decision-Making Under Uncertainty with Bayesian Networks
by: Jackson, Krystal, et al.
Published: (2026)
by: Jackson, Krystal, et al.
Published: (2026)
Intolerable Risk Threshold Recommendations for Artificial Intelligence
by: Raman, Deepika, et al.
Published: (2025)
by: Raman, Deepika, et al.
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models
by: Fan, Jigang, et al.
Published: (2025)
by: Fan, Jigang, et al.
Published: (2025)
A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems
by: Shaikh, Ayan Javeed, et al.
Published: (2026)
by: Shaikh, Ayan Javeed, et al.
Published: (2026)
Demo: ViolentUTF as An Accessible Platform for Generative AI Red Teaming
by: Nguyen, Tam n.
Published: (2025)
by: Nguyen, Tam n.
Published: (2025)
Red Teaming Methodology for Design Obfuscation
by: Liu, Yuntao, et al.
Published: (2025)
by: Liu, Yuntao, et al.
Published: (2025)
Autonomous Adversary: Red-Teaming in the age of LLM
by: Mamun, Mohammad, et al.
Published: (2026)
by: Mamun, Mohammad, et al.
Published: (2026)
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
by: Dheekonda, Raja Sekhar Rao, et al.
Published: (2026)
by: Dheekonda, Raja Sekhar Rao, et al.
Published: (2026)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
by: Sinha, Anusha, et al.
Published: (2025)
by: Sinha, Anusha, et al.
Published: (2025)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
by: Syros, Georgios, et al.
Published: (2026)
by: Syros, Georgios, et al.
Published: (2026)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing
by: Kaplan, Caelin, et al.
Published: (2025)
by: Kaplan, Caelin, et al.
Published: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Red Teaming Large Reasoning Models
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
by: Abuadbba, Alsharif, et al.
Published: (2025)
by: Abuadbba, Alsharif, et al.
Published: (2025)
Pwned: How Often Are Americans' Online Accounts Breached?
by: Cor, Ken, et al.
Published: (2018)
by: Cor, Ken, et al.
Published: (2018)
A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications
by: Srivastava, Shruti, et al.
Published: (2026)
by: Srivastava, Shruti, et al.
Published: (2026)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
by: Zhou, Andy, et al.
Published: (2025)
by: Zhou, Andy, et al.
Published: (2025)
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
by: Erez, Lidor, et al.
Published: (2026)
by: Erez, Lidor, et al.
Published: (2026)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
Red Team Redemption: A Structured Comparison of Open-Source Tools for Adversary Emulation
by: Landauer, Max, et al.
Published: (2024)
by: Landauer, Max, et al.
Published: (2024)
Benchmarking LLMs in an Embodied Environment for Blue Team Threat Hunting
by: Liu, Xiaoqun, et al.
Published: (2025)
by: Liu, Xiaoqun, et al.
Published: (2025)
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025)
by: Challita, Brian, et al.
Published: (2025)
LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems
by: Atta, Hammad, et al.
Published: (2026)
by: Atta, Hammad, et al.
Published: (2026)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
Beyond the Scope: Security Testing of Permission Management in Team Workspace
by: Wan, Liuhuo, et al.
Published: (2025)
by: Wan, Liuhuo, et al.
Published: (2025)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
by: Wang, Yanting, et al.
Published: (2026)
by: Wang, Yanting, et al.
Published: (2026)
Assessing the Trustworthiness of Electronic Identity Management Systems: Framework and Insights from Inception to Deployment
by: Bottarelli, Mirko, et al.
Published: (2025)
by: Bottarelli, Mirko, et al.
Published: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
by: Morasso, Cristian, et al.
Published: (2026)
by: Morasso, Cristian, et al.
Published: (2026)
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026)
by: Chia-Pei, et al.
Published: (2026)
ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation
by: Zhao, Haochen, et al.
Published: (2026)
by: Zhao, Haochen, et al.
Published: (2026)
Resource Consumption Red-Teaming for Large Vision-Language Models
by: Gao, Haoran, et al.
Published: (2025)
by: Gao, Haoran, et al.
Published: (2025)
A Red Teaming Roadmap Towards System-Level Safety
by: Wang, Zifan, et al.
Published: (2025)
by: Wang, Zifan, et al.
Published: (2025)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
Similar Items
-
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025) -
Toward Risk Thresholds for AI-Enabled Cyber Threats: Enhancing Decision-Making Under Uncertainty with Bayesian Networks
by: Jackson, Krystal, et al.
Published: (2026) -
Intolerable Risk Threshold Recommendations for Artificial Intelligence
by: Raman, Deepika, et al.
Published: (2025) -
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025) -
SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models
by: Fan, Jigang, et al.
Published: (2025)