Optimizing AI Agent Attacks With Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Loughridge, Chloe, Colognese, Paul, Griffin, Avery, Tracy, Tyler, Kutasov, Jon, Benton, Joe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918186439409664
author Loughridge, Chloe
Colognese, Paul
Griffin, Avery
Tracy, Tyler
Kutasov, Jon
Benton, Joe
author_facet Loughridge, Chloe
Colognese, Paul
Griffin, Avery
Tracy, Tyler
Kutasov, Jon
Benton, Joe
contents As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good control evaluations require eliciting strong attack policies. This can be challenging in complex agentic environments where compute constraints leave us data-poor. In this work, we show how to optimize attack policies in SHADE-Arena, a dataset of diverse realistic control environments. We do this by decomposing attack capability into five constituent skills -- suspicion modeling, attack selection, plan synthesis, execution and subtlety -- and optimizing each component individually. To get around the constraint of limited data, we develop a probabilistic model of attack dynamics, optimize our attack hyperparameters using this simulation, and then show that the results transfer to SHADE-Arena. This results in a substantial improvement in attack strength, reducing safety score from a baseline of 0.87 to 0.41 using our scaffold.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02823
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing AI Agent Attacks With Synthetic Data
Loughridge, Chloe
Colognese, Paul
Griffin, Avery
Tracy, Tyler
Kutasov, Jon
Benton, Joe
Artificial Intelligence
As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good control evaluations require eliciting strong attack policies. This can be challenging in complex agentic environments where compute constraints leave us data-poor. In this work, we show how to optimize attack policies in SHADE-Arena, a dataset of diverse realistic control environments. We do this by decomposing attack capability into five constituent skills -- suspicion modeling, attack selection, plan synthesis, execution and subtlety -- and optimizing each component individually. To get around the constraint of limited data, we develop a probabilistic model of attack dynamics, optimize our attack hyperparameters using this simulation, and then show that the results transfer to SHADE-Arena. This results in a substantial improvement in attack strength, reducing safety score from a baseline of 0.87 to 0.41 using our scaffold.
title Optimizing AI Agent Attacks With Synthetic Data
topic Artificial Intelligence
url https://arxiv.org/abs/2511.02823