DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Mukhopadhyay, Snehasis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI
von: Mukhopadhyay, Snehasis
Veröffentlicht: (2026)
von: Mukhopadhyay, Snehasis
Veröffentlicht: (2026)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
von: Lermen, Simon, et al.
Veröffentlicht: (2025)
von: Lermen, Simon, et al.
Veröffentlicht: (2025)
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026)
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026)
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
von: Chen, Kang, et al.
Veröffentlicht: (2024)
von: Chen, Kang, et al.
Veröffentlicht: (2024)
What if Deception Cannot be Detected? A Cross-Linguistic Study on the Limits of Deception Detection from Text
von: Velutharambath, Aswathy, et al.
Veröffentlicht: (2025)
von: Velutharambath, Aswathy, et al.
Veröffentlicht: (2025)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
Can Factual Statements be Deceptive? The DeFaBel Corpus of Belief-based Deception
von: Velutharambath, Aswathy, et al.
Veröffentlicht: (2024)
von: Velutharambath, Aswathy, et al.
Veröffentlicht: (2024)
Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
von: Zheng, Li, et al.
Veröffentlicht: (2026)
von: Zheng, Li, et al.
Veröffentlicht: (2026)
The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence
von: Wan, Herun, et al.
Veröffentlicht: (2026)
von: Wan, Herun, et al.
Veröffentlicht: (2026)
DECOR: Auditing LLM Deception via Information Manipulation Theory
von: Cai, Linyue, et al.
Veröffentlicht: (2026)
von: Cai, Linyue, et al.
Veröffentlicht: (2026)
Probing the Limits of the Lie Detector Approach to LLM Deception
von: Berger, Tom-Felix
Veröffentlicht: (2026)
von: Berger, Tom-Felix
Veröffentlicht: (2026)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2026)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2026)
How Entangled is Factuality and Deception in German?
von: Velutharambath, Aswathy, et al.
Veröffentlicht: (2024)
von: Velutharambath, Aswathy, et al.
Veröffentlicht: (2024)
Voting-based Multimodal Automatic Deception Detection
von: Touma, Lana, et al.
Veröffentlicht: (2023)
von: Touma, Lana, et al.
Veröffentlicht: (2023)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
An Assessment of Model-On-Model Deception
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
Detecting Deceptive Dark Patterns in E-commerce Platforms
von: Ramteke, Arya, et al.
Veröffentlicht: (2024)
von: Ramteke, Arya, et al.
Veröffentlicht: (2024)
AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models
von: Mukhopadhyay, Snehasis, et al.
Veröffentlicht: (2025)
von: Mukhopadhyay, Snehasis, et al.
Veröffentlicht: (2025)
Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings
von: Miah, Md Messal Monem, et al.
Veröffentlicht: (2025)
von: Miah, Md Messal Monem, et al.
Veröffentlicht: (2025)
Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL
von: Wongkamjan, Wichayaporn, et al.
Veröffentlicht: (2025)
von: Wongkamjan, Wichayaporn, et al.
Veröffentlicht: (2025)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
von: Rani, Anku, et al.
Veröffentlicht: (2023)
von: Rani, Anku, et al.
Veröffentlicht: (2023)
Exploring the Deceptive Power of LLM-Generated Fake News: A Study of Real-World Detection Challenges
von: Sun, Yanshen, et al.
Veröffentlicht: (2024)
von: Sun, Yanshen, et al.
Veröffentlicht: (2024)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
von: Sinha, Aarush, et al.
Veröffentlicht: (2026)
von: Sinha, Aarush, et al.
Veröffentlicht: (2026)
Deceptive Patterns of Intelligent and Interactive Writing Assistants
von: Benharrak, Karim, et al.
Veröffentlicht: (2024)
von: Benharrak, Karim, et al.
Veröffentlicht: (2024)
Effects of Soft-Domain Transfer and Named Entity Information on Deception Detection
von: Triplett, Steven, et al.
Veröffentlicht: (2024)
von: Triplett, Steven, et al.
Veröffentlicht: (2024)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
To Tell The Truth: Language of Deception and Language Models
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
From Deception to Detection: The Dual Roles of Large Language Models in Fake News
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
von: Kumar, Sachin
Veröffentlicht: (2026)
von: Kumar, Sachin
Veröffentlicht: (2026)
Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing
von: Xie, Jiakuan, et al.
Veröffentlicht: (2025)
von: Xie, Jiakuan, et al.
Veröffentlicht: (2025)
Semantic Deception: When Reasoning Models Can't Compute an Addition
von: de Leeuw, Nathaniël, et al.
Veröffentlicht: (2025)
von: de Leeuw, Nathaniël, et al.
Veröffentlicht: (2025)
Do Large Language Models Exhibit Spontaneous Rational Deception?
von: Taylor, Samuel M., et al.
Veröffentlicht: (2025)
von: Taylor, Samuel M., et al.
Veröffentlicht: (2025)
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
von: Bu, Yuyan, et al.
Veröffentlicht: (2026)
von: Bu, Yuyan, et al.
Veröffentlicht: (2026)
PU-Lie: Lightweight Deception Detection in Imbalanced Diplomatic Dialogues via Positive-Unlabeled Learning
von: Kuwar, Bhavinkumar Vinodbhai, et al.
Veröffentlicht: (2025)
von: Kuwar, Bhavinkumar Vinodbhai, et al.
Veröffentlicht: (2025)
Dharma, Data and Deception: An LLM-Powered Rhetorical Analysis of Cow-Urine Health Claims on YouTube
von: Munir, Sheza, et al.
Veröffentlicht: (2026)
von: Munir, Sheza, et al.
Veröffentlicht: (2026)
Humanlike Multi-user Agent (HUMA): Designing a Deceptively Human AI Facilitator for Group Chats
von: Jacniacki, Mateusz, et al.
Veröffentlicht: (2025)
von: Jacniacki, Mateusz, et al.
Veröffentlicht: (2025)
Domain-Independent Deception: A New Taxonomy and Linguistic Analysis
von: Verma, Rakesh M., et al.
Veröffentlicht: (2024)
von: Verma, Rakesh M., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI
von: Mukhopadhyay, Snehasis
Veröffentlicht: (2026) -
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
von: Lermen, Simon, et al.
Veröffentlicht: (2025) -
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
von: Xu, Yang, et al.
Veröffentlicht: (2025) -
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
von: Rassul, Yassin H., et al.
Veröffentlicht: (2026) -
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
von: Chen, Kang, et al.
Veröffentlicht: (2024)