SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Najt, Elle, Toft, Colin, Tracy, Tyler, Roger, Fabien, Benton, Joe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910234034831360
author Najt, Elle
Toft, Colin
Tracy, Tyler
Roger, Fabien
Benton, Joe
author_facet Najt, Elle
Toft, Colin
Tracy, Tyler
Roger, Fabien
Benton, Joe
contents Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dangerous misaligned behavior. To better understand the limitations of such monitors against the diverse attack strategies that a coding agent could use, we present SLEIGHT-Bench (Subtle Low-itEration Insight-Guided Harmful Transcripts), a benchmark of synthetic transcripts containing 40 attacks across 11 categories, each showing a coding agent covertly pursuing a harmful objective (e.g. weight exfiltration, credential theft, unmonitored rogue deployment). We find that 20 of 40 attacks are never caught by an Opus 4.6 monitor with extended thinking across 10 trials at a 1% false-positive rate, while the overall catch rate is 32%. We identify several evasion strategies such as exploiting model priors, user instruction ambiguity, and state manipulation, which facilitate creating attacks that evade frontier monitors. We also elicit stronger monitor performance using coding agents as monitors versus regular prompted monitors, and for some evasion strategies show improved catch rates with targeted monitor prompts. Our dataset and evaluation framework are available at https://github.com/safety-research/sleight-bench and https://huggingface.co/datasets/sleightbench/SLEIGHT-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16626
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Najt, Elle
Toft, Colin
Tracy, Tyler
Roger, Fabien
Benton, Joe
Cryptography and Security
Artificial Intelligence
Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dangerous misaligned behavior. To better understand the limitations of such monitors against the diverse attack strategies that a coding agent could use, we present SLEIGHT-Bench (Subtle Low-itEration Insight-Guided Harmful Transcripts), a benchmark of synthetic transcripts containing 40 attacks across 11 categories, each showing a coding agent covertly pursuing a harmful objective (e.g. weight exfiltration, credential theft, unmonitored rogue deployment). We find that 20 of 40 attacks are never caught by an Opus 4.6 monitor with extended thinking across 10 trials at a 1% false-positive rate, while the overall catch rate is 32%. We identify several evasion strategies such as exploiting model priors, user instruction ambiguity, and state manipulation, which facilitate creating attacks that evade frontier monitors. We also elicit stronger monitor performance using coding agents as monitors versus regular prompted monitors, and for some evasion strategies show improved catch rates with targeted monitor prompts. Our dataset and evaluation framework are available at https://github.com/safety-research/sleight-bench and https://huggingface.co/datasets/sleightbench/SLEIGHT-Bench.
title SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2605.16626