Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Meek, Austin, Sprejer, Eitan, Arcuschin, Iván, Brockmeier, Austin J., Basart, Steven
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918223503425536
author Meek, Austin
Sprejer, Eitan
Arcuschin, Iván
Brockmeier, Austin J.
Basart, Steven
author_facet Meek, Austin
Sprejer, Eitan
Arcuschin, Iván
Brockmeier, Austin J.
Basart, Steven
contents Chain-of-thought (CoT) outputs let us read a model's step-by-step reasoning. Since any long, serial reasoning process must pass through this textual trace, the quality of the CoT is a direct window into what the model is thinking. This visibility could help us spot unsafe or misaligned behavior (monitorability), but only if the CoT is transparent about its internal reasoning (faithfulness). Fully measuring faithfulness is difficult, so researchers often focus on examining the CoT in cases where the model changes its answer after adding a cue to the input. This proxy finds some instances of unfaithfulness but loses information when the model maintains its answer, and does not investigate aspects of reasoning not tied to the cue. We extend these results to a more holistic sense of monitorability by introducing verbosity: whether the CoT lists every factor needed to solve the task. We combine faithfulness and verbosity into a single monitorability score that shows how well the CoT serves as the model's external `working memory', a property that many safety schemes based on CoT monitoring depend on. We evaluate instruction-tuned and reasoning models on BBH, GPQA, and MMLU. Our results show that models can appear faithful yet remain hard to monitor when they leave out key factors, and that monitorability differs sharply across model families. We release our evaluation code using the Inspect library to support reproducible future work.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27378
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
Meek, Austin
Sprejer, Eitan
Arcuschin, Iván
Brockmeier, Austin J.
Basart, Steven
Machine Learning
Artificial Intelligence
Computation and Language
Chain-of-thought (CoT) outputs let us read a model's step-by-step reasoning. Since any long, serial reasoning process must pass through this textual trace, the quality of the CoT is a direct window into what the model is thinking. This visibility could help us spot unsafe or misaligned behavior (monitorability), but only if the CoT is transparent about its internal reasoning (faithfulness). Fully measuring faithfulness is difficult, so researchers often focus on examining the CoT in cases where the model changes its answer after adding a cue to the input. This proxy finds some instances of unfaithfulness but loses information when the model maintains its answer, and does not investigate aspects of reasoning not tied to the cue. We extend these results to a more holistic sense of monitorability by introducing verbosity: whether the CoT lists every factor needed to solve the task. We combine faithfulness and verbosity into a single monitorability score that shows how well the CoT serves as the model's external `working memory', a property that many safety schemes based on CoT monitoring depend on. We evaluate instruction-tuned and reasoning models on BBH, GPQA, and MMLU. Our results show that models can appear faithful yet remain hard to monitor when they leave out key factors, and that monitorability differs sharply across model families. We release our evaluation code using the Inspect library to support reproducible future work.
title Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.27378