DECOR: Auditing LLM Deception via Information Manipulation Theory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Linyue, Yeh, Samuel, Dhamala, Jwala, Gupta, Rahul, Li, Sharon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909055487836160
author Cai, Linyue
Yeh, Samuel
Dhamala, Jwala
Gupta, Rahul
Li, Sharon
author_facet Cai, Linyue
Yeh, Samuel
Dhamala, Jwala
Gupta, Rahul
Li, Sharon
contents Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detect. Existing black-box methods rely on coarse-grained judgments, offering limited interpretability and failing to pinpoint which facts were distorted and how. We introduce DECOR, a multi-agent framework grounded in Information Manipulation Theory for fine-grained auditing of strategic deception in LLM responses. DECOR decomposes input contexts into atomic informational units and scores each unit against the response across four dimensions of manipulation, producing interpretable manipulation profiles that are aggregated into a global deception index. We comprehensively evaluate DECOR on both single-turn and multi-turn deception detection benchmarks spanning real-world domains, and show that DECOR achieves state-of-the-art performance on both, outperforming competitive baselines. The framework generalizes across 15 frontier models, and ablation studies confirm the contribution of each key design component. Our findings demonstrate that fine-grained, theory-grounded auditing of information manipulation offers an effective and interpretable path for LLM deception detection.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19270
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DECOR: Auditing LLM Deception via Information Manipulation Theory
Cai, Linyue
Yeh, Samuel
Dhamala, Jwala
Gupta, Rahul
Li, Sharon
Computation and Language
Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detect. Existing black-box methods rely on coarse-grained judgments, offering limited interpretability and failing to pinpoint which facts were distorted and how. We introduce DECOR, a multi-agent framework grounded in Information Manipulation Theory for fine-grained auditing of strategic deception in LLM responses. DECOR decomposes input contexts into atomic informational units and scores each unit against the response across four dimensions of manipulation, producing interpretable manipulation profiles that are aggregated into a global deception index. We comprehensively evaluate DECOR on both single-turn and multi-turn deception detection benchmarks spanning real-world domains, and show that DECOR achieves state-of-the-art performance on both, outperforming competitive baselines. The framework generalizes across 15 frontier models, and ablation studies confirm the contribution of each key design component. Our findings demonstrate that fine-grained, theory-grounded auditing of information manipulation offers an effective and interpretable path for LLM deception detection.
title DECOR: Auditing LLM Deception via Information Manipulation Theory
topic Computation and Language
url https://arxiv.org/abs/2605.19270