LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Yang, Zhang, Xuanming, Yeh, Samuel, Dhamala, Jwala, Dia, Ousmane, Gupta, Rahul, Li, Sharon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DECOR: Auditing LLM Deception via Information Manipulation Theory
di: Cai, Linyue, et al.
Pubblicazione: (2026)
di: Cai, Linyue, et al.
Pubblicazione: (2026)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
di: Wu, Yichen, et al.
Pubblicazione: (2025)
di: Wu, Yichen, et al.
Pubblicazione: (2025)
Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
di: Zhang, Xuanming, et al.
Pubblicazione: (2025)
di: Zhang, Xuanming, et al.
Pubblicazione: (2025)
MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
di: Zhang, Xuanming, et al.
Pubblicazione: (2025)
di: Zhang, Xuanming, et al.
Pubblicazione: (2025)
DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents
di: Mukhopadhyay, Snehasis
Pubblicazione: (2026)
di: Mukhopadhyay, Snehasis
Pubblicazione: (2026)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
Deceptive Patterns of Intelligent and Interactive Writing Assistants
di: Benharrak, Karim, et al.
Pubblicazione: (2024)
di: Benharrak, Karim, et al.
Pubblicazione: (2024)
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
di: Chen, Kang, et al.
Pubblicazione: (2024)
di: Chen, Kang, et al.
Pubblicazione: (2024)
The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence
di: Wan, Herun, et al.
Pubblicazione: (2026)
di: Wan, Herun, et al.
Pubblicazione: (2026)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
di: Yeh, Samuel, et al.
Pubblicazione: (2025)
di: Yeh, Samuel, et al.
Pubblicazione: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
Can Factual Statements be Deceptive? The DeFaBel Corpus of Belief-based Deception
di: Velutharambath, Aswathy, et al.
Pubblicazione: (2024)
di: Velutharambath, Aswathy, et al.
Pubblicazione: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025)
di: Krishna, Satyapriya, et al.
Pubblicazione: (2025)
Do Large Language Models Exhibit Spontaneous Rational Deception?
di: Taylor, Samuel M., et al.
Pubblicazione: (2025)
di: Taylor, Samuel M., et al.
Pubblicazione: (2025)
Probing the Limits of the Lie Detector Approach to LLM Deception
di: Berger, Tom-Felix
Pubblicazione: (2026)
di: Berger, Tom-Felix
Pubblicazione: (2026)
What if Deception Cannot be Detected? A Cross-Linguistic Study on the Limits of Deception Detection from Text
di: Velutharambath, Aswathy, et al.
Pubblicazione: (2025)
di: Velutharambath, Aswathy, et al.
Pubblicazione: (2025)
How Entangled is Factuality and Deception in German?
di: Velutharambath, Aswathy, et al.
Pubblicazione: (2024)
di: Velutharambath, Aswathy, et al.
Pubblicazione: (2024)
An Assessment of Model-On-Model Deception
di: Heitkoetter, Julius, et al.
Pubblicazione: (2024)
di: Heitkoetter, Julius, et al.
Pubblicazione: (2024)
How Retrieved Context Shapes Internal Representations in RAG
di: Yeh, Samuel, et al.
Pubblicazione: (2026)
di: Yeh, Samuel, et al.
Pubblicazione: (2026)
Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
di: Zheng, Li, et al.
Pubblicazione: (2026)
di: Zheng, Li, et al.
Pubblicazione: (2026)
PU-Lie: Lightweight Deception Detection in Imbalanced Diplomatic Dialogues via Positive-Unlabeled Learning
di: Kuwar, Bhavinkumar Vinodbhai, et al.
Pubblicazione: (2025)
di: Kuwar, Bhavinkumar Vinodbhai, et al.
Pubblicazione: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
di: Rani, Anku, et al.
Pubblicazione: (2023)
di: Rani, Anku, et al.
Pubblicazione: (2023)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
To Tell The Truth: Language of Deception and Language Models
di: Hazra, Sanchaita, et al.
Pubblicazione: (2023)
di: Hazra, Sanchaita, et al.
Pubblicazione: (2023)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
di: Qian, Yusu, et al.
Pubblicazione: (2024)
di: Qian, Yusu, et al.
Pubblicazione: (2024)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
di: Kumar, Sachin
Pubblicazione: (2026)
di: Kumar, Sachin
Pubblicazione: (2026)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
di: Järviniemi, Olli, et al.
Pubblicazione: (2024)
di: Järviniemi, Olli, et al.
Pubblicazione: (2024)
Semantic Deception: When Reasoning Models Can't Compute an Addition
di: de Leeuw, Nathaniël, et al.
Pubblicazione: (2025)
di: de Leeuw, Nathaniël, et al.
Pubblicazione: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
di: Kang, Caixin, et al.
Pubblicazione: (2025)
di: Kang, Caixin, et al.
Pubblicazione: (2025)
Dharma, Data and Deception: An LLM-Powered Rhetorical Analysis of Cow-Urine Health Claims on YouTube
di: Munir, Sheza, et al.
Pubblicazione: (2026)
di: Munir, Sheza, et al.
Pubblicazione: (2026)
LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
di: Yeh, Samuel, et al.
Pubblicazione: (2025)
di: Yeh, Samuel, et al.
Pubblicazione: (2025)
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
di: Jaeger, Samuel, et al.
Pubblicazione: (2026)
di: Jaeger, Samuel, et al.
Pubblicazione: (2026)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
di: Ignat, Oana, et al.
Pubblicazione: (2024)
di: Ignat, Oana, et al.
Pubblicazione: (2024)
Deception Abilities Emerged in Large Language Models
di: Hagendorff, Thilo
Pubblicazione: (2023)
di: Hagendorff, Thilo
Pubblicazione: (2023)
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
Voting-based Multimodal Automatic Deception Detection
di: Touma, Lana, et al.
Pubblicazione: (2023)
di: Touma, Lana, et al.
Pubblicazione: (2023)
Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing
di: Xie, Jiakuan, et al.
Pubblicazione: (2025)
di: Xie, Jiakuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DECOR: Auditing LLM Deception via Information Manipulation Theory
di: Cai, Linyue, et al.
Pubblicazione: (2026) -
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
di: Wu, Yichen, et al.
Pubblicazione: (2025) -
Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
di: Zhang, Xuanming, et al.
Pubblicazione: (2025) -
MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
di: Zhang, Xuanming, et al.
Pubblicazione: (2025) -
DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents
di: Mukhopadhyay, Snehasis
Pubblicazione: (2026)