When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Kai, Zhang, Yihao, Sun, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913877860548608
author Wang, Kai
Zhang, Yihao
Sun, Meng
author_facet Wang, Kai
Zhang, Yihao
Sun, Meng
contents The honesty of large language models (LLMs) is a critical alignment challenge, especially as advanced systems with chain-of-thought (CoT) reasoning may strategically deceive humans. Unlike traditional honesty issues on LLMs, which could be possibly explained as some kind of hallucination, those models' explicit thought paths enable us to study strategic deception--goal-driven, intentional misinformation where reasoning contradicts outputs. Using representation engineering, we systematically induce, detect, and control such deception in CoT-enabled LLMs, extracting "deception vectors" via Linear Artificial Tomography (LAT) for 89% detection accuracy. Through activation steering, we achieve a 40% success rate in eliciting context-appropriate deception without explicit prompts, unveiling the specific honesty-related issue of reasoning models and providing tools for trustworthy AI alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
Wang, Kai
Zhang, Yihao
Sun, Meng
Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
The honesty of large language models (LLMs) is a critical alignment challenge, especially as advanced systems with chain-of-thought (CoT) reasoning may strategically deceive humans. Unlike traditional honesty issues on LLMs, which could be possibly explained as some kind of hallucination, those models' explicit thought paths enable us to study strategic deception--goal-driven, intentional misinformation where reasoning contradicts outputs. Using representation engineering, we systematically induce, detect, and control such deception in CoT-enabled LLMs, extracting "deception vectors" via Linear Artificial Tomography (LAT) for 89% detection accuracy. Through activation steering, we achieve a 40% success rate in eliciting context-appropriate deception without explicit prompts, unveiling the specific honesty-related issue of reasoning models and providing tools for trustworthy AI alignment.
title When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
topic Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.04909