Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: MacDermott, Matt, Wei, Qiyao, Djoneva, Rada, Ward, Francis Rhys
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918238769643520
author MacDermott, Matt
Wei, Qiyao
Djoneva, Rada
Ward, Francis Rhys
author_facet MacDermott, Matt
Wei, Qiyao
Djoneva, Rada
Ward, Francis Rhys
contents AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as the pursuit of harmful objectives. However, the extent to which CoT faithfully reflects the underlying reasoning process, and hence the extent to which it can be usefully monitored, may be influenced by certain aspects of training. We investigate how different \emph{training incentives}, applied to a reasoning model, affect its monitorability. We introduce a novel methodology for measuring monitorability according to whether a monitor can predict a key latent variable using the model's reasoning. When controlling for accuracy, we do not find evidence for consistent effects from commonly used incentives (length penalties and KL regularisation), but we find that adversarial optimisation (penalising monitor accuracy) degrades monitor performance, while direct optimisation for monitorability does not reliably lead to improvements. Our code is available at https://github.com/QiyaoWei/reasoning-under-pressure.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
MacDermott, Matt
Wei, Qiyao
Djoneva, Rada
Ward, Francis Rhys
Artificial Intelligence
Cryptography and Security
AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as the pursuit of harmful objectives. However, the extent to which CoT faithfully reflects the underlying reasoning process, and hence the extent to which it can be usefully monitored, may be influenced by certain aspects of training. We investigate how different \emph{training incentives}, applied to a reasoning model, affect its monitorability. We introduce a novel methodology for measuring monitorability according to whether a monitor can predict a key latent variable using the model's reasoning. When controlling for accuracy, we do not find evidence for consistent effects from commonly used incentives (length penalties and KL regularisation), but we find that adversarial optimisation (penalising monitor accuracy) degrades monitor performance, while direct optimisation for monitorability does not reliably lead to improvements. Our code is available at https://github.com/QiyaoWei/reasoning-under-pressure.
title Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2512.00218