Training a General Purpose Automated Red Teaming Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Padmakumar, Aishwarya, Derczynski, Leon, Rebedea, Traian, Parisien, Christopher
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911622445924352
author Padmakumar, Aishwarya
Derczynski, Leon
Rebedea, Traian
Parisien, Christopher
author_facet Padmakumar, Aishwarya
Derczynski, Leon
Rebedea, Traian
Parisien, Christopher
contents Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They can also adapt to each specific LLM to discover weaknesses unique to it. Most current automated red teaming methods are intended for tackling safety and content moderation. Thus, they make use of content safety models as evaluators and optimize for circumventing them, and as such, have not been tested with other adversarial intents not typically captured by these. We propose a pipeline for training a red teaming model that can generalize to arbitrary adversarial goals, including objectives it has not been directly trained on, and that does not depend on the existence of a pre-existing evaluator available at training time. We demonstrate that finetuning small models, such as Qwen3-8B, using this pipeline results in a substantial improvement in their ability to generate attacks for both in and out of domain adversarial goals.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23067
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training a General Purpose Automated Red Teaming Model
Padmakumar, Aishwarya
Derczynski, Leon
Rebedea, Traian
Parisien, Christopher
Cryptography and Security
Computation and Language
Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They can also adapt to each specific LLM to discover weaknesses unique to it. Most current automated red teaming methods are intended for tackling safety and content moderation. Thus, they make use of content safety models as evaluators and optimize for circumventing them, and as such, have not been tested with other adversarial intents not typically captured by these. We propose a pipeline for training a red teaming model that can generalize to arbitrary adversarial goals, including objectives it has not been directly trained on, and that does not depend on the existence of a pre-existing evaluator available at training time. We demonstrate that finetuning small models, such as Qwen3-8B, using this pipeline results in a substantial improvement in their ability to generate attacks for both in and out of domain adversarial goals.
title Training a General Purpose Automated Red Teaming Model
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2604.23067