CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arnav, Benjamin, Bernabeu-Pérez, Pablo, Helm-Burger, Nathan, Kostolansky, Tim, Whittingham, Hannes, Phuong, Mary
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917101710606336
author Arnav, Benjamin
Bernabeu-Pérez, Pablo
Helm-Burger, Nathan
Kostolansky, Tim
Whittingham, Hannes
Phuong, Mary
author_facet Arnav, Benjamin
Bernabeu-Pérez, Pablo
Helm-Burger, Nathan
Kostolansky, Tim
Whittingham, Hannes
Phuong, Mary
contents As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought (CoT) monitoring, wherein a weaker trusted monitor model continuously oversees the intermediate reasoning steps of a more powerful but untrusted model. We compare CoT monitoring to action-only monitoring, where only final outputs are reviewed, in a red-teaming setup where the untrusted model is instructed to pursue harmful side tasks while completing a coding problem. We find that while CoT monitoring is more effective than overseeing only model outputs in scenarios where action-only monitoring fails to reliably identify sabotage, reasoning traces can contain misleading rationalizations that deceive the CoT monitors, reducing performance in obvious sabotage cases. To address this, we introduce a hybrid protocol that independently scores model reasoning and actions, and combines them using a weighted average. Our hybrid monitor consistently outperforms both CoT and action-only monitors across all tested models and tasks, with detection rates twice higher than action-only monitoring for subtle deception scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23575
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Arnav, Benjamin
Bernabeu-Pérez, Pablo
Helm-Burger, Nathan
Kostolansky, Tim
Whittingham, Hannes
Phuong, Mary
Artificial Intelligence
Machine Learning
As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought (CoT) monitoring, wherein a weaker trusted monitor model continuously oversees the intermediate reasoning steps of a more powerful but untrusted model. We compare CoT monitoring to action-only monitoring, where only final outputs are reviewed, in a red-teaming setup where the untrusted model is instructed to pursue harmful side tasks while completing a coding problem. We find that while CoT monitoring is more effective than overseeing only model outputs in scenarios where action-only monitoring fails to reliably identify sabotage, reasoning traces can contain misleading rationalizations that deceive the CoT monitors, reducing performance in obvious sabotage cases. To address this, we introduce a hybrid protocol that independently scores model reasoning and actions, and combines them using a weighted average. Our hybrid monitor consistently outperforms both CoT and action-only monitors across all tested models and tasks, with detection rates twice higher than action-only monitoring for subtle deception scenarios.
title CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.23575