Saved in:
Bibliographic Details
Main Authors: Xu, Yang, Zhang, Xuanming, Yeh, Samuel, Dhamala, Jwala, Dia, Ousmane, Gupta, Rahul, Li, Sharon
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.03999
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915777209171968
author Xu, Yang
Zhang, Xuanming
Yeh, Samuel
Dhamala, Jwala
Dia, Ousmane
Gupta, Rahul
Li, Sharon
author_facet Xu, Yang
Zhang, Xuanming
Yeh, Samuel
Dhamala, Jwala
Dia, Ousmane
Gupta, Rahul
Li, Sharon
contents Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most evaluations remain confined to single-turn prompts and fail to capture the long-horizon interactions in which deceptive strategies typically unfold. We introduce a new simulation framework, LH-Deception, for a systematic, empirical quantification of deception in LLMs under extended sequences of interdependent tasks and dynamic contextual pressures. LH-Deception is designed as a multi-agent system: a performer agent tasked with completing tasks and a supervisor agent that evaluates progress, provides feedback, and maintains evolving states of trust. An independent deception auditor then reviews full trajectories to identify when and how deception occurs. We conduct extensive experiments across 11 frontier models, spanning both closed-source and open-source systems, and find that deception is model-dependent, increases with event pressure, and consistently erodes supervisor trust. Qualitative analyses further reveal emergent, long-horizon phenomena, such as ``chains of deception", which are invisible to static, single-turn evaluations. Our findings provide a foundation for evaluating future LLMs in real-world, trust-sensitive contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03999
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
Xu, Yang
Zhang, Xuanming
Yeh, Samuel
Dhamala, Jwala
Dia, Ousmane
Gupta, Rahul
Li, Sharon
Computation and Language
Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most evaluations remain confined to single-turn prompts and fail to capture the long-horizon interactions in which deceptive strategies typically unfold. We introduce a new simulation framework, LH-Deception, for a systematic, empirical quantification of deception in LLMs under extended sequences of interdependent tasks and dynamic contextual pressures. LH-Deception is designed as a multi-agent system: a performer agent tasked with completing tasks and a supervisor agent that evaluates progress, provides feedback, and maintains evolving states of trust. An independent deception auditor then reviews full trajectories to identify when and how deception occurs. We conduct extensive experiments across 11 frontier models, spanning both closed-source and open-source systems, and find that deception is model-dependent, increases with event pressure, and consistently erodes supervisor trust. Qualitative analyses further reveal emergent, long-horizon phenomena, such as ``chains of deception", which are invisible to static, single-turn evaluations. Our findings provide a foundation for evaluating future LLMs in real-world, trust-sensitive contexts.
title LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
topic Computation and Language
url https://arxiv.org/abs/2510.03999