DECOR: Auditing LLM Deception via Information Manipulation Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Linyue, Yeh, Samuel, Dhamala, Jwala, Gupta, Rahul, Li, Sharon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
by: Yeh, Samuel, et al.
Published: (2025)
by: Yeh, Samuel, et al.
Published: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
How Retrieved Context Shapes Internal Representations in RAG
by: Yeh, Samuel, et al.
Published: (2026)
by: Yeh, Samuel, et al.
Published: (2026)
LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
by: Yeh, Samuel, et al.
Published: (2025)
by: Yeh, Samuel, et al.
Published: (2025)
Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
by: Zhang, Xuanming, et al.
Published: (2025)
by: Zhang, Xuanming, et al.
Published: (2025)
MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
by: Zhang, Xuanming, et al.
Published: (2025)
by: Zhang, Xuanming, et al.
Published: (2025)
Auditing LLM Benchmarks with Item Response Theory
by: Land, Sander, et al.
Published: (2026)
by: Land, Sander, et al.
Published: (2026)
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
by: Sheshadri, Abhay, et al.
Published: (2026)
by: Sheshadri, Abhay, et al.
Published: (2026)
DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and Rewriting
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents
by: Mukhopadhyay, Snehasis
Published: (2026)
by: Mukhopadhyay, Snehasis
Published: (2026)
Smart Audit System Empowered by LLM
by: Yao, Xu, et al.
Published: (2024)
by: Yao, Xu, et al.
Published: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence
by: Wan, Herun, et al.
Published: (2026)
by: Wan, Herun, et al.
Published: (2026)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
PU-Lie: Lightweight Deception Detection in Imbalanced Diplomatic Dialogues via Positive-Unlabeled Learning
by: Kuwar, Bhavinkumar Vinodbhai, et al.
Published: (2025)
by: Kuwar, Bhavinkumar Vinodbhai, et al.
Published: (2025)
Do Large Language Models Exhibit Spontaneous Rational Deception?
by: Taylor, Samuel M., et al.
Published: (2025)
by: Taylor, Samuel M., et al.
Published: (2025)
Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education
by: Gupta, Amogh, et al.
Published: (2026)
by: Gupta, Amogh, et al.
Published: (2026)
CreAgentive: An Agent Workflow Driven Multi-Category Creative Generation Engine
by: Cheng, Yuyang, et al.
Published: (2025)
by: Cheng, Yuyang, et al.
Published: (2025)
Auditing Agent Harness Safety
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
Probing the Limits of the Lie Detector Approach to LLM Deception
by: Berger, Tom-Felix
Published: (2026)
by: Berger, Tom-Felix
Published: (2026)
Improving Narrative Classification and Explanation via Fine Tuned Language Models
by: Tyagi, Rishit, et al.
Published: (2025)
by: Tyagi, Rishit, et al.
Published: (2025)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
by: Wu, Yichen, et al.
Published: (2025)
by: Wu, Yichen, et al.
Published: (2025)
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
by: Li, Jiaqian, et al.
Published: (2026)
by: Li, Jiaqian, et al.
Published: (2026)
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
by: Chen, Kang, et al.
Published: (2024)
by: Chen, Kang, et al.
Published: (2024)
Semi-structured LLM Reasoners Can Be Rigorously Audited
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
SemanticShield: LLM-Powered Audits Expose Shilling Attacks in Recommender Systems
by: Li, Kaihong, et al.
Published: (2025)
by: Li, Kaihong, et al.
Published: (2025)
Can Factual Statements be Deceptive? The DeFaBel Corpus of Belief-based Deception
by: Velutharambath, Aswathy, et al.
Published: (2024)
by: Velutharambath, Aswathy, et al.
Published: (2024)
Semgrex and Ssurgeon, Searching and Manipulating Dependency Graphs
by: Bauer, John, et al.
Published: (2024)
by: Bauer, John, et al.
Published: (2024)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
by: Rani, Anku, et al.
Published: (2023)
by: Rani, Anku, et al.
Published: (2023)
Not What, But How: A Communicative Audit of LLM Response Framing
by: Pawar, Siddhesh Milind, et al.
Published: (2026)
by: Pawar, Siddhesh Milind, et al.
Published: (2026)
Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
by: Zheng, Li, et al.
Published: (2026)
by: Zheng, Li, et al.
Published: (2026)
What if Deception Cannot be Detected? A Cross-Linguistic Study on the Limits of Deception Detection from Text
by: Velutharambath, Aswathy, et al.
Published: (2025)
by: Velutharambath, Aswathy, et al.
Published: (2025)
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
by: Jaeger, Samuel, et al.
Published: (2026)
by: Jaeger, Samuel, et al.
Published: (2026)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
by: Rassul, Yassin H., et al.
Published: (2026)
by: Rassul, Yassin H., et al.
Published: (2026)
Dharma, Data and Deception: An LLM-Powered Rhetorical Analysis of Cow-Urine Health Claims on YouTube
by: Munir, Sheza, et al.
Published: (2026)
by: Munir, Sheza, et al.
Published: (2026)
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and Correction
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
How Entangled is Factuality and Deception in German?
by: Velutharambath, Aswathy, et al.
Published: (2024)
by: Velutharambath, Aswathy, et al.
Published: (2024)
Similar Items
-
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
by: Xu, Yang, et al.
Published: (2025) -
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
by: Yeh, Samuel, et al.
Published: (2025) -
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023) -
How Retrieved Context Shapes Internal Representations in RAG
by: Yeh, Samuel, et al.
Published: (2026) -
LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
by: Yeh, Samuel, et al.
Published: (2025)