RAPTOR: Ridge-Adaptive Logistic Probes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gao, Ziqi, Zhu, Yaotian, Zeng, Qingcheng, Zhao, Xu, Wang, Ziqing, Ruan, Feng, Ding, Kaize
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914304851181568
author Gao, Ziqi
Zhu, Yaotian
Zeng, Qingcheng
Zhao, Xu
Wang, Ziqing
Ruan, Feng
Ding, Kaize
author_facet Gao, Ziqi
Zhu, Yaotian
Zeng, Qingcheng
Zhao, Xu
Wang, Ziqing
Ruan, Feng
Ding, Kaize
contents Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted from a probe and injected via additive activation steering by adding it to a layer representation during the forward pass. The effectiveness of this pipeline hinges on estimating concept vectors that are accurate, directionally stable under ablation, and inexpensive to obtain. Motivated by these desiderata, we propose RAPTOR (Ridge-Adaptive Logistic Probe), a simple L2-regularized logistic probe whose validation-tuned ridge strength yields concept vectors from normalized weights. Across extensive experiments on instruction-tuned LLMs and human-written concept datasets, RAPTOR matches or exceeds strong baselines in accuracy while achieving competitive directional stability and substantially lower training cost; these quantitative results are supported by qualitative downstream steering demonstrations. Finally, using the Convex Gaussian Min-max Theorem (CGMT), we provide a mechanistic characterization of ridge logistic regression in an idealized Gaussian teacher-student model in the high-dimensional few-shot regime, explaining how penalty strength mediates probe accuracy and concept-vector stability and yielding structural predictions that qualitatively align with trends observed on real LLM embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00158
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RAPTOR: Ridge-Adaptive Logistic Probes
Gao, Ziqi
Zhu, Yaotian
Zeng, Qingcheng
Zhao, Xu
Wang, Ziqing
Ruan, Feng
Ding, Kaize
Machine Learning
Artificial Intelligence
Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted from a probe and injected via additive activation steering by adding it to a layer representation during the forward pass. The effectiveness of this pipeline hinges on estimating concept vectors that are accurate, directionally stable under ablation, and inexpensive to obtain. Motivated by these desiderata, we propose RAPTOR (Ridge-Adaptive Logistic Probe), a simple L2-regularized logistic probe whose validation-tuned ridge strength yields concept vectors from normalized weights. Across extensive experiments on instruction-tuned LLMs and human-written concept datasets, RAPTOR matches or exceeds strong baselines in accuracy while achieving competitive directional stability and substantially lower training cost; these quantitative results are supported by qualitative downstream steering demonstrations. Finally, using the Convex Gaussian Min-max Theorem (CGMT), we provide a mechanistic characterization of ridge logistic regression in an idealized Gaussian teacher-student model in the high-dimensional few-shot regime, explaining how penalty strength mediates probe accuracy and concept-vector stability and yielding structural predictions that qualitatively align with trends observed on real LLM embeddings.
title RAPTOR: Ridge-Adaptive Logistic Probes
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.00158