Mitigating Deceptive Alignment via Self-Monitoring

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Jiaming, Chen, Wenqi, Wang, Kaile, Hong, Donghai, Fang, Sitong, Chen, Boyuan, Zhou, Jiayi, Dai, Juntao, Han, Sirui, Guo, Yike, Yang, Yaodong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913857677557760
author Ji, Jiaming
Chen, Wenqi
Wang, Kaile
Hong, Donghai
Fang, Sitong
Chen, Boyuan
Zhou, Jiayi
Dai, Juntao
Han, Sirui
Guo, Yike
Yang, Yaodong
author_facet Ji, Jiaming
Chen, Wenqi
Wang, Kaile
Hong, Donghai
Fang, Sitong
Chen, Boyuan
Zhou, Jiayi
Dai, Juntao
Han, Sirui
Guo, Yike
Yang, Yaodong
contents Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which a model appears aligned while covertly pursuing misaligned goals. Existing safety pipelines treat deception as a black-box output to be filtered post-hoc, leaving the model free to scheme during its internal reasoning. We ask: Can deception be intercepted while the model is thinking? We answer this question, the first framework that embeds a Self-Monitor inside the CoT process itself, named CoT Monitor+. During generation, the model produces (i) ordinary reasoning steps and (ii) an internal self-evaluation signal trained to flag and suppress misaligned strategies. The signal is used as an auxiliary reward in reinforcement learning, creating a feedback loop that rewards honest reasoning and discourages hidden goals. To study deceptive alignment systematically, we introduce DeceptionBench, a five-category benchmark that probes covert alignment-faking, sycophancy, etc. We evaluate various LLMs and show that unrestricted CoT roughly aggravates the deceptive tendency. In contrast, CoT Monitor+ cuts deceptive behaviors by 43.8% on average while preserving task accuracy. Further, when the self-monitor signal replaces an external weak judge in RL fine-tuning, models exhibit substantially fewer obfuscated thoughts and retain transparency. Our project website can be found at cot-monitor-plus.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2505_18807
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Deceptive Alignment via Self-Monitoring
Ji, Jiaming
Chen, Wenqi
Wang, Kaile
Hong, Donghai
Fang, Sitong
Chen, Boyuan
Zhou, Jiayi
Dai, Juntao
Han, Sirui
Guo, Yike
Yang, Yaodong
Artificial Intelligence
Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which a model appears aligned while covertly pursuing misaligned goals. Existing safety pipelines treat deception as a black-box output to be filtered post-hoc, leaving the model free to scheme during its internal reasoning. We ask: Can deception be intercepted while the model is thinking? We answer this question, the first framework that embeds a Self-Monitor inside the CoT process itself, named CoT Monitor+. During generation, the model produces (i) ordinary reasoning steps and (ii) an internal self-evaluation signal trained to flag and suppress misaligned strategies. The signal is used as an auxiliary reward in reinforcement learning, creating a feedback loop that rewards honest reasoning and discourages hidden goals. To study deceptive alignment systematically, we introduce DeceptionBench, a five-category benchmark that probes covert alignment-faking, sycophancy, etc. We evaluate various LLMs and show that unrestricted CoT roughly aggravates the deceptive tendency. In contrast, CoT Monitor+ cuts deceptive behaviors by 43.8% on average while preserving task accuracy. Further, when the self-monitor signal replaces an external weak judge in RL fine-tuning, models exhibit substantially fewer obfuscated thoughts and retain transparency. Our project website can be found at cot-monitor-plus.github.io
title Mitigating Deceptive Alignment via Self-Monitoring
topic Artificial Intelligence
url https://arxiv.org/abs/2505.18807