Linear probes rely on textual evidence: Results from leakage mitigation studies in language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boxo, Gerard, Neelappa, Aman, Raval, Shivam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915847073693696
author Boxo, Gerard
Neelappa, Aman
Raval, Shivam
author_facet Boxo, Gerard
Neelappa, Aman
Raval, Shivam
contents White-box monitors are a popular technique for detecting potentially harmful behaviours in language models. While they perform well in general, their effectiveness in detecting text-ambiguous behaviour is disputed. In this work, we find evidence that removing textual evidence of a behaviour significantly decreases probe performance. The AUROC reduction ranges from $10$- to $30$-point depending on the setting. We evaluate probe monitors across three setups (Sandbagging, Sycophancy, and Bias), finding that when probes rely on textual evidence of the target behaviour (such as system prompts or CoT reasoning), performance degrades once these tokens are filtered. This filtering procedure is standard practice for output monitor evaluation. As further evidence of this phenomenon, we train Model Organisms which produce outputs without any behaviour verbalisations. We validate that probe performance on Model Organisms is substantially lower than unfiltered evaluations: $0.57$ vs $0.74$ AUROC for Bias, and $0.57$ vs $0.94$ AUROC for Sandbagging. Our findings suggest that linear probes may be brittle in scenarios where they must detect non-surface-level patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
Boxo, Gerard
Neelappa, Aman
Raval, Shivam
Artificial Intelligence
Computation and Language
Machine Learning
White-box monitors are a popular technique for detecting potentially harmful behaviours in language models. While they perform well in general, their effectiveness in detecting text-ambiguous behaviour is disputed. In this work, we find evidence that removing textual evidence of a behaviour significantly decreases probe performance. The AUROC reduction ranges from $10$- to $30$-point depending on the setting. We evaluate probe monitors across three setups (Sandbagging, Sycophancy, and Bias), finding that when probes rely on textual evidence of the target behaviour (such as system prompts or CoT reasoning), performance degrades once these tokens are filtered. This filtering procedure is standard practice for output monitor evaluation. As further evidence of this phenomenon, we train Model Organisms which produce outputs without any behaviour verbalisations. We validate that probe performance on Model Organisms is substantially lower than unfiltered evaluations: $0.57$ vs $0.74$ AUROC for Bias, and $0.57$ vs $0.94$ AUROC for Sandbagging. Our findings suggest that linear probes may be brittle in scenarios where they must detect non-surface-level patterns.
title Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.21344