Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.03999 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915777209171968 |
|---|---|
| author | Xu, Yang Zhang, Xuanming Yeh, Samuel Dhamala, Jwala Dia, Ousmane Gupta, Rahul Li, Sharon |
| author_facet | Xu, Yang Zhang, Xuanming Yeh, Samuel Dhamala, Jwala Dia, Ousmane Gupta, Rahul Li, Sharon |
| contents | Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most evaluations remain confined to single-turn prompts and fail to capture the long-horizon interactions in which deceptive strategies typically unfold. We introduce a new simulation framework, LH-Deception, for a systematic, empirical quantification of deception in LLMs under extended sequences of interdependent tasks and dynamic contextual pressures. LH-Deception is designed as a multi-agent system: a performer agent tasked with completing tasks and a supervisor agent that evaluates progress, provides feedback, and maintains evolving states of trust. An independent deception auditor then reviews full trajectories to identify when and how deception occurs. We conduct extensive experiments across 11 frontier models, spanning both closed-source and open-source systems, and find that deception is model-dependent, increases with event pressure, and consistently erodes supervisor trust. Qualitative analyses further reveal emergent, long-horizon phenomena, such as ``chains of deception", which are invisible to static, single-turn evaluations. Our findings provide a foundation for evaluating future LLMs in real-world, trust-sensitive contexts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03999 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions Xu, Yang Zhang, Xuanming Yeh, Samuel Dhamala, Jwala Dia, Ousmane Gupta, Rahul Li, Sharon Computation and Language Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most evaluations remain confined to single-turn prompts and fail to capture the long-horizon interactions in which deceptive strategies typically unfold. We introduce a new simulation framework, LH-Deception, for a systematic, empirical quantification of deception in LLMs under extended sequences of interdependent tasks and dynamic contextual pressures. LH-Deception is designed as a multi-agent system: a performer agent tasked with completing tasks and a supervisor agent that evaluates progress, provides feedback, and maintains evolving states of trust. An independent deception auditor then reviews full trajectories to identify when and how deception occurs. We conduct extensive experiments across 11 frontier models, spanning both closed-source and open-source systems, and find that deception is model-dependent, increases with event pressure, and consistently erodes supervisor trust. Qualitative analyses further reveal emergent, long-horizon phenomena, such as ``chains of deception", which are invisible to static, single-turn evaluations. Our findings provide a foundation for evaluating future LLMs in real-world, trust-sensitive contexts. |
| title | LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.03999 |