Reasoning Models Don't Always Say What They Think

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Yanda, Benton, Joe, Radhakrishnan, Ansh, Uesato, Jonathan, Denison, Carson, Schulman, John, Somani, Arushi, Hase, Peter, Wagner, Misha, Roger, Fabien, Mikulik, Vlad, Bowman, Samuel R., Leike, Jan, Kaplan, Jared, Perez, Ethan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912366835269632
author Chen, Yanda
Benton, Joe
Radhakrishnan, Ansh
Uesato, Jonathan
Denison, Carson
Schulman, John
Somani, Arushi
Hase, Peter
Wagner, Misha
Roger, Fabien
Mikulik, Vlad
Bowman, Samuel R.
Leike, Jan
Kaplan, Jared
Perez, Ethan
author_facet Chen, Yanda
Benton, Joe
Radhakrishnan, Ansh
Uesato, Jonathan
Denison, Carson
Schulman, John
Somani, Arushi
Hase, Peter
Wagner, Misha
Roger, Fabien
Mikulik, Vlad
Bowman, Samuel R.
Leike, Jan
Kaplan, Jared
Perez, Ethan
contents Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effectiveness of such monitoring hinges on CoTs faithfully representing models' actual reasoning processes. We evaluate CoT faithfulness of state-of-the-art reasoning models across 6 reasoning hints presented in the prompts and find: (1) for most settings and models tested, CoTs reveal their usage of hints in at least 1% of examples where they use the hint, but the reveal rate is often below 20%, (2) outcome-based reinforcement learning initially improves faithfulness but plateaus without saturating, and (3) when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor. These results suggest that CoT monitoring is a promising way of noticing undesired behaviors during training and evaluations, but that it is not sufficient to rule them out. They also suggest that in settings like ours where CoT reasoning is not necessary, test-time monitoring of CoTs is unlikely to reliably catch rare and catastrophic unexpected behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning Models Don't Always Say What They Think
Chen, Yanda
Benton, Joe
Radhakrishnan, Ansh
Uesato, Jonathan
Denison, Carson
Schulman, John
Somani, Arushi
Hase, Peter
Wagner, Misha
Roger, Fabien
Mikulik, Vlad
Bowman, Samuel R.
Leike, Jan
Kaplan, Jared
Perez, Ethan
Computation and Language
Artificial Intelligence
Machine Learning
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effectiveness of such monitoring hinges on CoTs faithfully representing models' actual reasoning processes. We evaluate CoT faithfulness of state-of-the-art reasoning models across 6 reasoning hints presented in the prompts and find: (1) for most settings and models tested, CoTs reveal their usage of hints in at least 1% of examples where they use the hint, but the reveal rate is often below 20%, (2) outcome-based reinforcement learning initially improves faithfulness but plateaus without saturating, and (3) when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor. These results suggest that CoT monitoring is a promising way of noticing undesired behaviors during training and evaluations, but that it is not sufficient to rule them out. They also suggest that in settings like ours where CoT reasoning is not necessary, test-time monitoring of CoTs is unlikely to reliably catch rare and catastrophic unexpected behaviors.
title Reasoning Models Don't Always Say What They Think
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.05410