Detecting Strategic Deception Using Linear Probes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goldowsky-Dill, Nicholas, Chughtai, Bilal, Heimersheim, Stefan, Hobbhahn, Marius |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Difficulties with Evaluating a Deception Detector for AIs
von: Smith, Lewis, et al.
Veröffentlicht: (2025)
von: Smith, Lewis, et al.
Veröffentlicht: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
von: Braun, Dan, et al.
Veröffentlicht: (2024)
von: Braun, Dan, et al.
Veröffentlicht: (2024)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
You can remove GPT2's LayerNorm by fine-tuning
von: Heimersheim, Stefan
Veröffentlicht: (2024)
von: Heimersheim, Stefan
Veröffentlicht: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
How to use and interpret activation patching
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
Training on Documents About Monitoring Leads to CoT Obfuscation
von: Haskins, Reilly, et al.
Veröffentlicht: (2026)
von: Haskins, Reilly, et al.
Veröffentlicht: (2026)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
Can Language Models Explain Their Own Classification Behavior?
von: Sherburn, Dane, et al.
Veröffentlicht: (2024)
von: Sherburn, Dane, et al.
Veröffentlicht: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
Technical Report: Evaluating Goal Drift in Language Model Agents
von: Arike, Rauno, et al.
Veröffentlicht: (2025)
von: Arike, Rauno, et al.
Veröffentlicht: (2025)
Building Production-Ready Probes For Gemini
von: Kramár, János, et al.
Veröffentlicht: (2026)
von: Kramár, János, et al.
Veröffentlicht: (2026)
Strategically Deceptive Model Deployment in Performative Prediction
von: Bautiste, Javier Sanguino, et al.
Veröffentlicht: (2025)
von: Bautiste, Javier Sanguino, et al.
Veröffentlicht: (2025)
Bi-GRU Based Deception Detection using EEG Signals
von: Avola, Danilo, et al.
Veröffentlicht: (2025)
von: Avola, Danilo, et al.
Veröffentlicht: (2025)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
Building Better Deception Probes Using Targeted Instruction Pairs
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
von: Daniels, Oliver, et al.
Veröffentlicht: (2026)
von: Daniels, Oliver, et al.
Veröffentlicht: (2026)
Robust Filtering -- Novel Statistical Learning and Inference Algorithms with Applications
von: Chughtai, Aamir Hussain
Veröffentlicht: (2025)
von: Chughtai, Aamir Hussain
Veröffentlicht: (2025)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
von: Braun, Dan, et al.
Veröffentlicht: (2025)
von: Braun, Dan, et al.
Veröffentlicht: (2025)
Evolution of SAE Features Across Layers in LLMs
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
Frontier Models are Capable of In-context Scheming
von: Meinke, Alexander, et al.
Veröffentlicht: (2024)
von: Meinke, Alexander, et al.
Veröffentlicht: (2024)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Deception Detection: From Static Texts to Multimodal Signals
von: Logan, Mandela
Veröffentlicht: (2025)
von: Logan, Mandela
Veröffentlicht: (2025)
Probing the Limits of the Lie Detector Approach to LLM Deception
von: Berger, Tom-Felix
Veröffentlicht: (2026)
von: Berger, Tom-Felix
Veröffentlicht: (2026)
Training Deliberative Monitors for Black-Box Scheming Detection
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
Characterizing stable regions in the residual stream of LLMs
von: Janiak, Jett, et al.
Veröffentlicht: (2024)
von: Janiak, Jett, et al.
Veröffentlicht: (2024)
Towards evaluations-based safety cases for AI scheming
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
Deception Detection in Dyadic Exchanges Using Multimodal Machine Learning: A Study on a Swedish Cohort
von: Samuels, Thomas Jack, et al.
Veröffentlicht: (2025)
von: Samuels, Thomas Jack, et al.
Veröffentlicht: (2025)
Strategic Classification with Non-Linear Classifiers
von: Trachtenberg, Benyamin, et al.
Veröffentlicht: (2025)
von: Trachtenberg, Benyamin, et al.
Veröffentlicht: (2025)
Linear Strategic Classification with Endogenous Improvements
von: Shrivastava, Siddharth, et al.
Veröffentlicht: (2026)
von: Shrivastava, Siddharth, et al.
Veröffentlicht: (2026)
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
von: Stix, Charlotte, et al.
Veröffentlicht: (2025)
von: Stix, Charlotte, et al.
Veröffentlicht: (2025)
Flexible inference in heterogeneous and attributed multilayer networks
von: Contisciani, Martina, et al.
Veröffentlicht: (2024)
von: Contisciani, Martina, et al.
Veröffentlicht: (2024)
Evolving and Detecting Multi-Turn Deception using Geometric Signatures
von: Kumar, Surender Suresh, et al.
Veröffentlicht: (2026)
von: Kumar, Surender Suresh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024) -
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024) -
Difficulties with Evaluating a Deception Detector for AIs
von: Smith, Lewis, et al.
Veröffentlicht: (2025) -
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026) -
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)