Prompt Optimization and Evaluation for LLM Automated Red Teaming

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Freenor, Michael, Alvarez, Lauren, Leal, Milton, Smith, Lily, Garrett, Joel, Husieva, Yelyzaveta, Woodruff, Madeline, Miller, Ryan, Kummerfeld, Erich, Medeiros, Rafael, Schulhoff, Sander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909711365832704
author Freenor, Michael
Alvarez, Lauren
Leal, Milton
Smith, Lily
Garrett, Joel
Husieva, Yelyzaveta
Woodruff, Madeline
Miller, Ryan
Kummerfeld, Erich
Medeiros, Rafael
Schulhoff, Sander
author_facet Freenor, Michael
Alvarez, Lauren
Leal, Milton
Smith, Lily
Garrett, Joel
Husieva, Yelyzaveta
Woodruff, Madeline
Miller, Ryan
Kummerfeld, Erich
Medeiros, Rafael
Schulhoff, Sander
contents Applications that use Large Language Models (LLMs) are becoming widespread, making the identification of system vulnerabilities increasingly important. Automated Red Teaming accelerates this effort by using an LLM to generate and execute attacks against target systems. Attack generators are evaluated using the Attack Success Rate (ASR) the sample mean calculated over the judgment of success for each attack. In this paper, we introduce a method for optimizing attack generator prompts that applies ASR to individual attacks. By repeating each attack multiple times against a randomly seeded target, we measure an attack's discoverability the expectation of the individual attack success. This approach reveals exploitable patterns that inform prompt optimization, ultimately enabling more robust evaluation and refinement of generators.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt Optimization and Evaluation for LLM Automated Red Teaming
Freenor, Michael
Alvarez, Lauren
Leal, Milton
Smith, Lily
Garrett, Joel
Husieva, Yelyzaveta
Woodruff, Madeline
Miller, Ryan
Kummerfeld, Erich
Medeiros, Rafael
Schulhoff, Sander
Cryptography and Security
Computation and Language
Applications that use Large Language Models (LLMs) are becoming widespread, making the identification of system vulnerabilities increasingly important. Automated Red Teaming accelerates this effort by using an LLM to generate and execute attacks against target systems. Attack generators are evaluated using the Attack Success Rate (ASR) the sample mean calculated over the judgment of success for each attack. In this paper, we introduce a method for optimizing attack generator prompts that applies ASR to individual attacks. By repeating each attack multiple times against a randomly seeded target, we measure an attack's discoverability the expectation of the individual attack success. This approach reveals exploitable patterns that inform prompt optimization, ultimately enabling more robust evaluation and refinement of generators.
title Prompt Optimization and Evaluation for LLM Automated Red Teaming
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2507.22133