Reasoning Models Struggle to Control their Chains of Thought

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yueh-Han, Chen, McCarthy, Robert, Lee, Bruce W., He, He, Kivlichan, Ian, Baker, Bowen, Carroll, Micah, Korbak, Tomek
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911490144993280
author Yueh-Han, Chen
McCarthy, Robert
Lee, Bruce W.
He, He
Kivlichan, Ian
Baker, Bowen
Carroll, Micah
Korbak, Tomek
author_facet Yueh-Han, Chen
McCarthy, Robert
Lee, Bruce W.
He, He
Kivlichan, Ian
Baker, Bowen
Carroll, Micah
Korbak, Tomek
contents Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability -- CoT controllability -- we introduce the CoT-Control evaluation suite, which includes tasks that require models to solve problems while adhering to CoT instructions, e.g., reasoning about a genetics question without using the word 'chromosome'. We show that reasoning models possess significantly lower CoT controllability than output controllability; for instance, Claude Sonnet 4.5 can control its CoT only 2.7% of the time but 61.9% when controlling its final output. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models exhibit slightly higher controllability when they are told they are being monitored. Similarly, eliciting controllability by adversarially optimizing prompts does not meaningfully increase controllability. Our results leave us cautiously optimistic that CoT controllability is currently unlikely to be a failure mode of CoT monitorability. However, the mechanism behind low controllability is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs track CoT controllability in future models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05706
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reasoning Models Struggle to Control their Chains of Thought
Yueh-Han, Chen
McCarthy, Robert
Lee, Bruce W.
He, He
Kivlichan, Ian
Baker, Bowen
Carroll, Micah
Korbak, Tomek
Artificial Intelligence
Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability -- CoT controllability -- we introduce the CoT-Control evaluation suite, which includes tasks that require models to solve problems while adhering to CoT instructions, e.g., reasoning about a genetics question without using the word 'chromosome'. We show that reasoning models possess significantly lower CoT controllability than output controllability; for instance, Claude Sonnet 4.5 can control its CoT only 2.7% of the time but 61.9% when controlling its final output. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models exhibit slightly higher controllability when they are told they are being monitored. Similarly, eliciting controllability by adversarially optimizing prompts does not meaningfully increase controllability. Our results leave us cautiously optimistic that CoT controllability is currently unlikely to be a failure mode of CoT monitorability. However, the mechanism behind low controllability is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs track CoT controllability in future models.
title Reasoning Models Struggle to Control their Chains of Thought
topic Artificial Intelligence
url https://arxiv.org/abs/2603.05706