Probing the Limits of the Lie Detector Approach to LLM Deception
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Berger, Tom-Felix |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)
von: Levinstein, B. A., et al.
Veröffentlicht: (2023)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
von: Wang, Kai, et al.
Veröffentlicht: (2025)
von: Wang, Kai, et al.
Veröffentlicht: (2025)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
von: Kumar, Sachin
Veröffentlicht: (2026)
von: Kumar, Sachin
Veröffentlicht: (2026)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
On Limitations of LLM as Annotator for Low Resource Languages
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)
Effects of Soft-Domain Transfer and Named Entity Information on Deception Detection
von: Triplett, Steven, et al.
Veröffentlicht: (2024)
von: Triplett, Steven, et al.
Veröffentlicht: (2024)
The Zero Body Problem: Probing LLM Use of Sensory Language
von: Hicke, Rebecca M. M., et al.
Veröffentlicht: (2025)
von: Hicke, Rebecca M. M., et al.
Veröffentlicht: (2025)
LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
von: Schouten, Stefan F., et al.
Veröffentlicht: (2025)
von: Schouten, Stefan F., et al.
Veröffentlicht: (2025)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
LLM Flow Processes for Text-Conditioned Regression
von: Biggs, Felix, et al.
Veröffentlicht: (2026)
von: Biggs, Felix, et al.
Veröffentlicht: (2026)
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
von: Thede, Lukas, et al.
Veröffentlicht: (2025)
von: Thede, Lukas, et al.
Veröffentlicht: (2025)
Improving LLM Final Representations with Inter-Layer Geometry
von: Ulanovski, Tom, et al.
Veröffentlicht: (2026)
von: Ulanovski, Tom, et al.
Veröffentlicht: (2026)
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023)
von: Hagendorff, Thilo
Veröffentlicht: (2023)
More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
von: White, Colin, et al.
Veröffentlicht: (2024)
von: White, Colin, et al.
Veröffentlicht: (2024)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2024)
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2024)
Enhancing Robustness in Biomedical NLI Models: A Probing Approach for Clinical Trials
von: Mustafa, Ata
Veröffentlicht: (2024)
von: Mustafa, Ata
Veröffentlicht: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024)
von: Kramár, János, et al.
Veröffentlicht: (2024)
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
von: Qu, Jiaming, et al.
Veröffentlicht: (2025)
von: Qu, Jiaming, et al.
Veröffentlicht: (2025)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Dialectical Behavior Therapy Approach to LLM Prompting
von: Vitman, Oxana, et al.
Veröffentlicht: (2024)
von: Vitman, Oxana, et al.
Veröffentlicht: (2024)
Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
von: Mamtani, Sumit, et al.
Veröffentlicht: (2025)
von: Mamtani, Sumit, et al.
Veröffentlicht: (2025)
RAFT: Realistic Attacks to Fool Text Detectors
von: Wang, James, et al.
Veröffentlicht: (2024)
von: Wang, James, et al.
Veröffentlicht: (2024)
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
von: Wang, Sophie L., et al.
Veröffentlicht: (2026)
von: Wang, Sophie L., et al.
Veröffentlicht: (2026)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
von: Sharma, Akshat, et al.
Veröffentlicht: (2024)
von: Sharma, Akshat, et al.
Veröffentlicht: (2024)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
von: Jiang, Junqi, et al.
Veröffentlicht: (2025)
von: Jiang, Junqi, et al.
Veröffentlicht: (2025)
Rhetorical Questions in LLM Representations: A Linear Probing Study
von: Yao, Louie Hong, et al.
Veröffentlicht: (2026)
von: Yao, Louie Hong, et al.
Veröffentlicht: (2026)
LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text
von: Mihaila, George, et al.
Veröffentlicht: (2026)
von: Mihaila, George, et al.
Veröffentlicht: (2026)
Probing the Emergence of Cross-lingual Alignment during LLM Training
von: Wang, Hetong, et al.
Veröffentlicht: (2024)
von: Wang, Hetong, et al.
Veröffentlicht: (2024)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
von: Järviniemi, Olli, et al.
Veröffentlicht: (2024)
von: Järviniemi, Olli, et al.
Veröffentlicht: (2024)
Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
von: Bronzini, Marco, et al.
Veröffentlicht: (2025)
von: Bronzini, Marco, et al.
Veröffentlicht: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2024)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2024)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
von: Schoenegger, Loris, et al.
Veröffentlicht: (2024)
von: Schoenegger, Loris, et al.
Veröffentlicht: (2024)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Graph-Convolutional Autoencoder Ensembles for the Humanities, Illustrated with a Study of the American Slave Trade
von: Lippincott, Tom
Veröffentlicht: (2024)
von: Lippincott, Tom
Veröffentlicht: (2024)
Ähnliche Einträge
-
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
von: Levinstein, B. A., et al.
Veröffentlicht: (2023) -
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
von: Wang, Kai, et al.
Veröffentlicht: (2025) -
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
von: Kumar, Sachin
Veröffentlicht: (2026) -
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
von: Huang, Yao, et al.
Veröffentlicht: (2025) -
On Limitations of LLM as Annotator for Low Resource Languages
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)