Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Zhaomin, Du, Mingzhe, Ng, See-Kiong, He, Bingsheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909005451886592
author Wu, Zhaomin
Du, Mingzhe
Ng, See-Kiong
He, Bingsheng
author_facet Wu, Zhaomin
Du, Mingzhe
Ng, See-Kiong
He, Bingsheng
contents Large Language Models (LLMs) are widely deployed in reasoning, planning, and decision-making tasks, making their trustworthiness critical. A significant and underexplored risk is intentional deception, where an LLM deliberately fabricates or conceals information to serve a hidden objective. Existing studies typically induce deception by explicitly setting a hidden objective through prompting or fine-tuning, which may not reflect real-world human-LLM interactions. Moving beyond such human-induced deception, we investigate LLMs' self-initiated deception on benign prompts. To address the absence of ground truth, we propose a framework based on Contact Searching Questions (CSQ). This framework introduces two statistical metrics derived from psychological principles to quantify the likelihood of deception. The first, the Deceptive Intention Score, measures the model's bias toward a hidden objective. The second, the Deceptive Behavior Score, measures the inconsistency between the LLM's internal belief and its expressed output. Evaluating 16 leading LLMs, we find that both metrics rise in parallel and escalate with task difficulty for most models. Moreover, increasing model capacity does not always reduce deception, posing a significant challenge for future LLM development.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
Wu, Zhaomin
Du, Mingzhe
Ng, See-Kiong
He, Bingsheng
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) are widely deployed in reasoning, planning, and decision-making tasks, making their trustworthiness critical. A significant and underexplored risk is intentional deception, where an LLM deliberately fabricates or conceals information to serve a hidden objective. Existing studies typically induce deception by explicitly setting a hidden objective through prompting or fine-tuning, which may not reflect real-world human-LLM interactions. Moving beyond such human-induced deception, we investigate LLMs' self-initiated deception on benign prompts. To address the absence of ground truth, we propose a framework based on Contact Searching Questions (CSQ). This framework introduces two statistical metrics derived from psychological principles to quantify the likelihood of deception. The first, the Deceptive Intention Score, measures the model's bias toward a hidden objective. The second, the Deceptive Behavior Score, measures the inconsistency between the LLM's internal belief and its expressed output. Evaluating 16 leading LLMs, we find that both metrics rise in parallel and escalate with task difficulty for most models. Moreover, increasing model capacity does not always reduce deception, posing a significant challenge for future LLM development.
title Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.06361