Training a General Purpose Automated Red Teaming Model
Fuente:
arXiv
Saved in:
| Main Authors: | Padmakumar, Aishwarya, Derczynski, Leon, Rebedea, Traian, Parisien, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
by: Inie, Nanna, et al.
Published: (2023)
by: Inie, Nanna, et al.
Published: (2023)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
garak: A Framework for Security Probing Large Language Models
by: Derczynski, Leon, et al.
Published: (2024)
by: Derczynski, Leon, et al.
Published: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
by: Wei, Zhang, et al.
Published: (2025)
by: Wei, Zhang, et al.
Published: (2025)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
by: Wang, Zilong, et al.
Published: (2025)
by: Wang, Zilong, et al.
Published: (2025)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
by: Krishna, Arjun, et al.
Published: (2025)
by: Krishna, Arjun, et al.
Published: (2025)
Resource Consumption Red-Teaming for Large Vision-Language Models
by: Gao, Haoran, et al.
Published: (2025)
by: Gao, Haoran, et al.
Published: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
by: Zhang, Jinchuan, et al.
Published: (2024)
by: Zhang, Jinchuan, et al.
Published: (2024)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
Unsupervised Extraction of Dialogue Policies from Conversations
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
by: Xu, Huiyu, et al.
Published: (2024)
by: Xu, Huiyu, et al.
Published: (2024)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
by: Pathade, Chetan
Published: (2025)
by: Pathade, Chetan
Published: (2025)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024)
by: Kumar, Anurakt, et al.
Published: (2024)
Towards Inference-time Category-wise Safety Steering for Large Language Models
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
by: Verma, Apurv, et al.
Published: (2024)
by: Verma, Apurv, et al.
Published: (2024)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
by: Liu, Yanjiang, et al.
Published: (2025)
by: Liu, Yanjiang, et al.
Published: (2025)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025)
by: Nakash, Itay, et al.
Published: (2025)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
by: Borkar, Jaydeep, et al.
Published: (2025)
by: Borkar, Jaydeep, et al.
Published: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)
by: Chen, Shuo, et al.
Published: (2024)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
by: Yuan, Xiaohan, et al.
Published: (2024)
by: Yuan, Xiaohan, et al.
Published: (2024)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
by: Dou, Yipu, et al.
Published: (2026)
by: Dou, Yipu, et al.
Published: (2026)
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026)
by: Huang, Lige, et al.
Published: (2026)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
by: Zhao, Andrew, et al.
Published: (2024)
by: Zhao, Andrew, et al.
Published: (2024)
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)
by: Xu, Chejian, et al.
Published: (2024)
Contextual Agent Security: A Policy for Every Purpose
by: Tsai, Lillian, et al.
Published: (2025)
by: Tsai, Lillian, et al.
Published: (2025)
Similar Items
-
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
by: Inie, Nanna, et al.
Published: (2023) -
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024) -
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025) -
garak: A Framework for Security Probing Large Language Models
by: Derczynski, Leon, et al.
Published: (2024) -
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
by: Wei, Zhang, et al.
Published: (2025)