Saved in:
Bibliographic Details
Main Authors: Martin, Sam, Roger, Fabien
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.12366
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914559761055744
author Martin, Sam
Roger, Fabien
author_facet Martin, Sam
Roger, Fabien
contents Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely contain transcripts longer than 100K tokens. We show that when used as classifiers, current frontier models fail to notice dangerous actions more often in longer transcripts. In particular, on a dataset that requires identifying when a coding agent takes a subtly dangerous action, Opus 4.6, GPT 5.4, and Gemini 3.1 miss these actions $2\times$ to $30\times$ more often when they occur after 800K tokens of benign activity than when they occur on their own. We also show that these weaknesses can be partially mitigated with prompting techniques such as periodic reminders throughout the transcript and may be mitigated further with better post-training. Monitor evaluations that do not consider long-context degradation are likely overestimating monitor performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Classifier Context Rot: Monitor Performance Degrades with Context Length
Martin, Sam
Roger, Fabien
Artificial Intelligence
Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely contain transcripts longer than 100K tokens. We show that when used as classifiers, current frontier models fail to notice dangerous actions more often in longer transcripts. In particular, on a dataset that requires identifying when a coding agent takes a subtly dangerous action, Opus 4.6, GPT 5.4, and Gemini 3.1 miss these actions $2\times$ to $30\times$ more often when they occur after 800K tokens of benign activity than when they occur on their own. We also show that these weaknesses can be partially mitigated with prompting techniques such as periodic reminders throughout the transcript and may be mitigated further with better post-training. Monitor evaluations that do not consider long-context degradation are likely overestimating monitor performance.
title Classifier Context Rot: Monitor Performance Degrades with Context Length
topic Artificial Intelligence
url https://arxiv.org/abs/2605.12366