Benchmarking Deception Probes via Black-to-White Performance Boosts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Parrack, Avi, Attubato, Carlo Leonardo, Heimersheim, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914259624001536
author Parrack, Avi
Attubato, Carlo Leonardo
Heimersheim, Stefan
author_facet Parrack, Avi
Attubato, Carlo Leonardo
Heimersheim, Stefan
contents AI assistants will occasionally respond deceptively to user queries. Recently, linear classifiers (called "deception probes") have been trained to distinguish the internal activations of a language model during deceptive versus honest responses. However, it's unclear how effective these probes are at detecting deception in practice, nor whether such probes are resistant to simple counter strategies from a deceptive assistant who wishes to evade detection. In this paper, we compare white-box monitoring (where the monitor has access to token-level probe activations) to black-box monitoring (without such access). We benchmark deception probes by the extent to which the white box monitor outperforms the black-box monitor, i.e. the black-to-white performance boost. We find weak but encouraging black-to-white performance boosts from existing deception probes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Deception Probes via Black-to-White Performance Boosts
Parrack, Avi
Attubato, Carlo Leonardo
Heimersheim, Stefan
Artificial Intelligence
Machine Learning
68T01
I.2.7; K.4.1
AI assistants will occasionally respond deceptively to user queries. Recently, linear classifiers (called "deception probes") have been trained to distinguish the internal activations of a language model during deceptive versus honest responses. However, it's unclear how effective these probes are at detecting deception in practice, nor whether such probes are resistant to simple counter strategies from a deceptive assistant who wishes to evade detection. In this paper, we compare white-box monitoring (where the monitor has access to token-level probe activations) to black-box monitoring (without such access). We benchmark deception probes by the extent to which the white box monitor outperforms the black-box monitor, i.e. the black-to-white performance boost. We find weak but encouraging black-to-white performance boosts from existing deception probes.
title Benchmarking Deception Probes via Black-to-White Performance Boosts
topic Artificial Intelligence
Machine Learning
68T01
I.2.7; K.4.1
url https://arxiv.org/abs/2507.12691