Building Better Deception Probes Using Targeted Instruction Pairs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Natarajan, Vikram, Jain, Devina, Arora, Shivam, Golechha, Satvik, Bloom, Joseph
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908803254976512
author Natarajan, Vikram
Jain, Devina
Arora, Shivam
Golechha, Satvik
Bloom, Joseph
author_facet Natarajan, Vikram
Jain, Devina
Arora, Shivam
Golechha, Satvik
Bloom, Joseph
contents Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforward scenarios, including spurious correlations and false positives on non-deceptive responses. In this paper, we identify the importance of the instruction pair used during training. Furthermore, we show that targeting specific deceptive behaviors through a human-interpretable taxonomy of deception leads to improved results on evaluation datasets. Our findings reveal that instruction pairs capture deceptive intent rather than content-specific patterns, explaining why prompt choice dominates probe performance (70.6% of variance). Given the heterogeneity of deception types across datasets, we conclude that organizations should design specialized probes targeting their specific threat models rather than seeking a universal deception detector.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01425
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Building Better Deception Probes Using Targeted Instruction Pairs
Natarajan, Vikram
Jain, Devina
Arora, Shivam
Golechha, Satvik
Bloom, Joseph
Artificial Intelligence
Machine Learning
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforward scenarios, including spurious correlations and false positives on non-deceptive responses. In this paper, we identify the importance of the instruction pair used during training. Furthermore, we show that targeting specific deceptive behaviors through a human-interpretable taxonomy of deception leads to improved results on evaluation datasets. Our findings reveal that instruction pairs capture deceptive intent rather than content-specific patterns, explaining why prompt choice dominates probe performance (70.6% of variance). Given the heterogeneity of deception types across datasets, we conclude that organizations should design specialized probes targeting their specific threat models rather than seeking a universal deception detector.
title Building Better Deception Probes Using Targeted Instruction Pairs
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.01425