Difficulties with Evaluating a Deception Detector for AIs
Fuente:
arXiv
Saved in:
| Main Authors: | Smith, Lewis, Chughtai, Bilal, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
Detecting Strategic Deception Using Linear Probes
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026)
by: Kramár, János, et al.
Published: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
by: Makelov, Aleksandar, et al.
Published: (2024)
by: Makelov, Aleksandar, et al.
Published: (2024)
Training on Documents About Monitoring Leads to CoT Obfuscation
by: Haskins, Reilly, et al.
Published: (2026)
by: Haskins, Reilly, et al.
Published: (2026)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
How to use and interpret activation patching
by: Heimersheim, Stefan, et al.
Published: (2024)
by: Heimersheim, Stefan, et al.
Published: (2024)
Can Language Models Explain Their Own Classification Behavior?
by: Sherburn, Dane, et al.
Published: (2024)
by: Sherburn, Dane, et al.
Published: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Explorations of Self-Repair in Language Models
by: Rushing, Cody, et al.
Published: (2024)
by: Rushing, Cody, et al.
Published: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
by: Zhang, Fred, et al.
Published: (2023)
by: Zhang, Fred, et al.
Published: (2023)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
by: Minder, Julian, et al.
Published: (2025)
by: Minder, Julian, et al.
Published: (2025)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
BatchTopK Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2024)
by: Bussmann, Bart, et al.
Published: (2024)
Transcoders Find Interpretable LLM Feature Circuits
by: Dunefsky, Jacob, et al.
Published: (2024)
by: Dunefsky, Jacob, et al.
Published: (2024)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Reasoning-Finetuning Repurposes Latent Representations in Base Models
by: Ward, Jake, et al.
Published: (2025)
by: Ward, Jake, et al.
Published: (2025)
Probing the Limits of the Lie Detector Approach to LLM Deception
by: Berger, Tom-Felix
Published: (2026)
by: Berger, Tom-Felix
Published: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Robust Filtering -- Novel Statistical Learning and Inference Algorithms with Applications
by: Chughtai, Aamir Hussain
Published: (2025)
by: Chughtai, Aamir Hussain
Published: (2025)
Can Go AIs be adversarially robust?
by: Tseng, Tom, et al.
Published: (2024)
by: Tseng, Tom, et al.
Published: (2024)
RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
by: Jafari, Farnoush Rezaei, et al.
Published: (2025)
by: Jafari, Farnoush Rezaei, et al.
Published: (2025)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2025)
by: Bussmann, Bart, et al.
Published: (2025)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
by: Kramár, János, et al.
Published: (2024)
by: Kramár, János, et al.
Published: (2024)
Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Bi-GRU Based Deception Detection using EEG Signals
by: Avola, Danilo, et al.
Published: (2025)
by: Avola, Danilo, et al.
Published: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
by: Ferrando, Javier, et al.
Published: (2024)
by: Ferrando, Javier, et al.
Published: (2024)
Interpreting Attention Layer Outputs with Sparse Autoencoders
by: Kissane, Connor, et al.
Published: (2024)
by: Kissane, Connor, et al.
Published: (2024)
Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
by: Majumdar, Anirudha
Published: (2025)
by: Majumdar, Anirudha
Published: (2025)
Model Organisms for Emergent Misalignment
by: Turner, Edward, et al.
Published: (2025)
by: Turner, Edward, et al.
Published: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Simple LLM Baselines are Competitive for Model Diffing
by: Kempf, Elias, et al.
Published: (2026)
by: Kempf, Elias, et al.
Published: (2026)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
by: Macar, Uzay, et al.
Published: (2025)
by: Macar, Uzay, et al.
Published: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
by: Bogdan, Paul C., et al.
Published: (2025)
by: Bogdan, Paul C., et al.
Published: (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
by: Wang, Atticus, et al.
Published: (2025)
by: Wang, Atticus, et al.
Published: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
Similar Items
-
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024) -
Detecting Strategic Deception Using Linear Probes
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025) -
Building Production-Ready Probes For Gemini
by: Kramár, János, et al.
Published: (2026) -
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
by: Makelov, Aleksandar, et al.
Published: (2024) -
Training on Documents About Monitoring Leads to CoT Obfuscation
by: Haskins, Reilly, et al.
Published: (2026)