MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Han, Sun, Yifan, Ko, Brian, Talati, Mann, Gong, Jiawen, Li, Zimeng, Yu, Naicheng, Yu, Xucheng, Shen, Wei, Jolly, Vedant, Zhang, Huan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914439121338368
author Wang, Han
Sun, Yifan
Ko, Brian
Talati, Mann
Gong, Jiawen
Li, Zimeng
Yu, Naicheng
Yu, Xucheng
Shen, Wei
Jolly, Vedant
Zhang, Huan
author_facet Wang, Han
Sun, Yifan
Ko, Brian
Talati, Mann
Gong, Jiawen
Li, Zimeng
Yu, Naicheng
Yu, Xucheng
Shen, Wei
Jolly, Vedant
Zhang, Huan
contents Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer faithfully reflects the actual reasons (i.e., decision-critical factors) driving the model's behavior, leading to the reduced CoT monitorability problem. However, a comprehensive and fully open-source benchmark for thoroughly evaluating CoT monitorability remains lacking. To address this gap, we propose MonitorBench, a systematic benchmark for evaluating CoT monitorability in LLMs. MonitorBench provides: (1) a diverse set of 1,514 test instances with carefully designed decision-critical factors across 19 tasks spanning 7 categories to characterize \textit{when} CoTs can be used to monitor the factors driving LLM behavior; and (2) two stress-test settings to quantify \textit{the extent to which} CoT monitorability can be degraded. Extensive experiments across multiple popular LLMs with varying capabilities show that CoT monitorability is higher when the decision-critical factors shape the intermediate reasoning process without merely influencing the final answer. More capable LLMs tend to exhibit lower monitorability. And all evaluated LLMs can intentionally reduce monitorability under stress-tests, with monitorability dropping by up to 30\% in some tasks that do not require structural reasoning over the decision-critical factors. Overall, MonitorBench provides a basis for further research on evaluating future LLMs, studying advanced stress-test monitorability techniques, and developing new monitoring approaches. The code is available at https://github.com/ASTRAL-Group/MonitorBench.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28590
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
Wang, Han
Sun, Yifan
Ko, Brian
Talati, Mann
Gong, Jiawen
Li, Zimeng
Yu, Naicheng
Yu, Xucheng
Shen, Wei
Jolly, Vedant
Zhang, Huan
Artificial Intelligence
Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer faithfully reflects the actual reasons (i.e., decision-critical factors) driving the model's behavior, leading to the reduced CoT monitorability problem. However, a comprehensive and fully open-source benchmark for thoroughly evaluating CoT monitorability remains lacking. To address this gap, we propose MonitorBench, a systematic benchmark for evaluating CoT monitorability in LLMs. MonitorBench provides: (1) a diverse set of 1,514 test instances with carefully designed decision-critical factors across 19 tasks spanning 7 categories to characterize \textit{when} CoTs can be used to monitor the factors driving LLM behavior; and (2) two stress-test settings to quantify \textit{the extent to which} CoT monitorability can be degraded. Extensive experiments across multiple popular LLMs with varying capabilities show that CoT monitorability is higher when the decision-critical factors shape the intermediate reasoning process without merely influencing the final answer. More capable LLMs tend to exhibit lower monitorability. And all evaluated LLMs can intentionally reduce monitorability under stress-tests, with monitorability dropping by up to 30\% in some tasks that do not require structural reasoning over the decision-critical factors. Overall, MonitorBench provides a basis for further research on evaluating future LLMs, studying advanced stress-test monitorability techniques, and developing new monitoring approaches. The code is available at https://github.com/ASTRAL-Group/MonitorBench.
title MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2603.28590