No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cencerrado, Iván Vicente Moreno, Masdemont, Arnau Padrés, Hawthorne, Anton Gonzalvez, Africa, David Demitri, Pacchiardi, Lorenzo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908860839624704
author Cencerrado, Iván Vicente Moreno
Masdemont, Arnau Padrés
Hawthorne, Anton Gonzalvez
Africa, David Demitri
Pacchiardi, Lorenzo
author_facet Cencerrado, Iván Vicente Moreno
Masdemont, Arnau Padrés
Hawthorne, Anton Gonzalvez
Africa, David Demitri
Pacchiardi, Lorenzo
contents Do large language models (LLMs) anticipate when they will answer correctly? To study this, we extract activations after a question is read but before any tokens are generated, and train linear probes to predict whether the model's forthcoming answer will be correct. Across three open-source model families ranging from 7 to 70 billion parameters, projections on this "in-advance correctness direction" trained on generic trivia questions predict success in distribution and on diverse out-of-distribution knowledge datasets, indicating a deeper signal than dataset-specific spurious features, and outperforming black-box baselines and verbalised predicted confidence. Predictive power saturates in intermediate layers and, notably, generalisation falters on questions requiring mathematical reasoning. Moreover, for models responding "I don't know", doing so strongly correlates with the probe score, indicating that the same direction also captures confidence. By complementing previous results on truthfulness and other behaviours obtained with probes and sparse auto-encoders, our work contributes essential findings to elucidate LLM internals.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
Cencerrado, Iván Vicente Moreno
Masdemont, Arnau Padrés
Hawthorne, Anton Gonzalvez
Africa, David Demitri
Pacchiardi, Lorenzo
Computation and Language
Artificial Intelligence
Do large language models (LLMs) anticipate when they will answer correctly? To study this, we extract activations after a question is read but before any tokens are generated, and train linear probes to predict whether the model's forthcoming answer will be correct. Across three open-source model families ranging from 7 to 70 billion parameters, projections on this "in-advance correctness direction" trained on generic trivia questions predict success in distribution and on diverse out-of-distribution knowledge datasets, indicating a deeper signal than dataset-specific spurious features, and outperforming black-box baselines and verbalised predicted confidence. Predictive power saturates in intermediate layers and, notably, generalisation falters on questions requiring mathematical reasoning. Moreover, for models responding "I don't know", doing so strongly correlates with the probe score, indicating that the same direction also captures confidence. By complementing previous results on truthfulness and other behaviours obtained with probes and sparse auto-encoders, our work contributes essential findings to elucidate LLM internals.
title No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.10625