Red Teaming AI Red Teaming
Fuente:
arXiv
Saved in:
| Main Authors: | Majumdar, Subhabrata, Pendleton, Brian, Gupta, Abhishek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025)
by: Challita, Brian, et al.
Published: (2025)
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025)
by: Ahmad, Lama, et al.
Published: (2025)
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
by: Kennedy, Wm. Matthew, et al.
Published: (2025)
by: Kennedy, Wm. Matthew, et al.
Published: (2025)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
by: Sinha, Anusha, et al.
Published: (2025)
by: Sinha, Anusha, et al.
Published: (2025)
Red Teaming Large Reasoning Models
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
by: Zhou, Andy, et al.
Published: (2025)
by: Zhou, Andy, et al.
Published: (2025)
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
by: Dheekonda, Raja Sekhar Rao, et al.
Published: (2026)
by: Dheekonda, Raja Sekhar Rao, et al.
Published: (2026)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
by: Xu, Huiyu, et al.
Published: (2024)
by: Xu, Huiyu, et al.
Published: (2024)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
by: Debi, Tanusree, et al.
Published: (2026)
by: Debi, Tanusree, et al.
Published: (2026)
BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing
by: Kaplan, Caelin, et al.
Published: (2025)
by: Kaplan, Caelin, et al.
Published: (2025)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
by: Singer, Brian, et al.
Published: (2025)
by: Singer, Brian, et al.
Published: (2025)
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025)
by: Nakash, Itay, et al.
Published: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
by: Quaye, Jessica, et al.
Published: (2024)
by: Quaye, Jessica, et al.
Published: (2024)
A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications
by: Srivastava, Shruti, et al.
Published: (2026)
by: Srivastava, Shruti, et al.
Published: (2026)
A Red Teaming Roadmap Towards System-Level Safety
by: Wang, Zifan, et al.
Published: (2025)
by: Wang, Zifan, et al.
Published: (2025)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024)
by: Kumar, Anurakt, et al.
Published: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
by: Liu, Yanjiang, et al.
Published: (2025)
by: Liu, Yanjiang, et al.
Published: (2025)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
by: Zhou, Zhaojiacheng
Published: (2026)
by: Zhou, Zhaojiacheng
Published: (2026)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
by: Syros, Georgios, et al.
Published: (2026)
by: Syros, Georgios, et al.
Published: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026)
by: Chia-Pei, et al.
Published: (2026)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
by: Ou, Haoran, et al.
Published: (2025)
by: Ou, Haoran, et al.
Published: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
by: Du, Pengfei
Published: (2025)
by: Du, Pengfei
Published: (2025)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
by: Yao, Hongwei, et al.
Published: (2026)
by: Yao, Hongwei, et al.
Published: (2026)
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
by: Bhatt, Manish, et al.
Published: (2025)
by: Bhatt, Manish, et al.
Published: (2025)
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
by: Rawat, Ambrish, et al.
Published: (2024)
by: Rawat, Ambrish, et al.
Published: (2024)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
by: Xie, Yuchong, et al.
Published: (2025)
by: Xie, Yuchong, et al.
Published: (2025)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
by: Mukherjee, Kunal, et al.
Published: (2026)
by: Mukherjee, Kunal, et al.
Published: (2026)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
by: Min, Nay Myat, et al.
Published: (2026)
by: Min, Nay Myat, et al.
Published: (2026)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
Similar Items
-
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025) -
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025) -
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
by: Kennedy, Wm. Matthew, et al.
Published: (2025) -
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024) -
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
by: Sinha, Anusha, et al.
Published: (2025)