Prompt Optimization and Evaluation for LLM Automated Red Teaming
Fuente:
arXiv
Saved in:
| Main Authors: | Freenor, Michael, Alvarez, Lauren, Leal, Milton, Smith, Lily, Garrett, Joel, Husieva, Yelyzaveta, Woodruff, Madeline, Miller, Ryan, Kummerfeld, Erich, Medeiros, Rafael, Schulhoff, Sander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mapping Semantic & Syntactic Relationships with Geometric Rotation
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
Better Simulations for Validating Causal Discovery with the DAG-Adaptation of the Onion Method
by: Andrews, Bryan, et al.
Published: (2024)
by: Andrews, Bryan, et al.
Published: (2024)
Temporal discontinuity trials and randomization: success rates versus design strength
by: Knaeble, Brian, et al.
Published: (2024)
by: Knaeble, Brian, et al.
Published: (2024)
Disordering the Establishment
by: Woodruff, Lily
Published: (2020)
by: Woodruff, Lily
Published: (2020)
An extensive simulation study evaluating the interaction of resampling techniques across multiple causal discovery contexts
by: Banerjee, Ritwick, et al.
Published: (2025)
by: Banerjee, Ritwick, et al.
Published: (2025)
In Response to Ecological Momentary Assessment of Voice and Psychological Factors: Group and Individual Mechanisms
by: Stephanie Misono, et al.
Published: (2025)
by: Stephanie Misono, et al.
Published: (2025)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
by: Deng, Wesley Hanwen, et al.
Published: (2025)
by: Deng, Wesley Hanwen, et al.
Published: (2025)
AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
by: Liu, Yufan, et al.
Published: (2025)
by: Liu, Yufan, et al.
Published: (2025)
Sensitivity Analysis of the Consistency Assumption
by: Knaeble, Brian, et al.
Published: (2025)
by: Knaeble, Brian, et al.
Published: (2025)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
by: Hardy, Amelia F., et al.
Published: (2024)
by: Hardy, Amelia F., et al.
Published: (2024)
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs
by: Mo, Wenjie Jacky, et al.
Published: (2025)
by: Mo, Wenjie Jacky, et al.
Published: (2025)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
by: Han, Vernon Toh Yan, et al.
Published: (2024)
by: Han, Vernon Toh Yan, et al.
Published: (2024)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)
by: Cuevas, Alejandro, et al.
Published: (2025)
GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and Doves
by: Peskoff, Denis, et al.
Published: (2024)
by: Peskoff, Denis, et al.
Published: (2024)
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2026)
by: Deng, Wesley Hanwen, et al.
Published: (2026)
Data for the general resistance of a 50-layer rectangular circuit with Rx=1Ω and Ry=xΩ
by: Kryvonis, Yelyzaveta, et al.
Published: (2024)
by: Kryvonis, Yelyzaveta, et al.
Published: (2024)
Comb Tensor Networks vs. Matrix Product States: Enhanced Efficiency in High-Dimensional Spaces
by: Kolesnyk, Danylo, et al.
Published: (2024)
by: Kolesnyk, Danylo, et al.
Published: (2024)
Optimal transient growth and transition to turbulence in the MHD pipe flow subject to a transverse magnetic field
by: Velizhanina, Yelyzaveta, et al.
Published: (2024)
by: Velizhanina, Yelyzaveta, et al.
Published: (2024)
Dynamic Phase Transitions in Mean-Field Ginzburg-Landau Models: Conjugate Fields and Fourier-Mode Scaling
by: Satynska, Yelyzaveta, et al.
Published: (2025)
by: Satynska, Yelyzaveta, et al.
Published: (2025)
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
by: Pavlova, Maya, et al.
Published: (2024)
by: Pavlova, Maya, et al.
Published: (2024)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
by: Singhania, Abhishek, et al.
Published: (2025)
by: Singhania, Abhishek, et al.
Published: (2025)
Effective Automation to Support the Human Infrastructure in AI Red Teaming
by: Zhang, Alice Qian, et al.
Published: (2025)
by: Zhang, Alice Qian, et al.
Published: (2025)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
by: Jung, MinJae, et al.
Published: (2026)
by: Jung, MinJae, et al.
Published: (2026)
Examining Causal Pathways to Suicidal Ideation and Nonsuicidal Self‐Injury in the Adolescent Brain Cognitive Development Study
by: Marvin Yan, et al.
Published: (2025)
by: Marvin Yan, et al.
Published: (2025)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
by: Yin, Chenlong, et al.
Published: (2026)
by: Yin, Chenlong, et al.
Published: (2026)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
by: Chin, Zhi-Yi, et al.
Published: (2023)
by: Chin, Zhi-Yi, et al.
Published: (2023)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
by: Schulhoff, Sander, et al.
Published: (2023)
by: Schulhoff, Sander, et al.
Published: (2023)
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
by: Morasso, Cristian, et al.
Published: (2026)
by: Morasso, Cristian, et al.
Published: (2026)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
by: Yu, Jiahao, et al.
Published: (2023)
by: Yu, Jiahao, et al.
Published: (2023)
EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection
by: Lu, Yijie, et al.
Published: (2025)
by: Lu, Yijie, et al.
Published: (2025)
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
Similar Items
-
Mapping Semantic & Syntactic Relationships with Geometric Rotation
by: Freenor, Michael, et al.
Published: (2025) -
Better Simulations for Validating Causal Discovery with the DAG-Adaptation of the Onion Method
by: Andrews, Bryan, et al.
Published: (2024) -
Temporal discontinuity trials and randomization: success rates versus design strength
by: Knaeble, Brian, et al.
Published: (2024) -
Disordering the Establishment
by: Woodruff, Lily
Published: (2020) -
An extensive simulation study evaluating the interaction of resampling techniques across multiple causal discovery contexts
by: Banerjee, Ritwick, et al.
Published: (2025)