What do your logits know? (The answer may surprise you!)

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fedzechkina, Masha, Gualdoni, Eleonora, Ramos, Rita, Williamson, Sinead
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913022286495744
author Fedzechkina, Masha
Gualdoni, Eleonora
Ramos, Rita
Williamson, Sinead
author_facet Fedzechkina, Masha
Gualdoni, Eleonora
Ramos, Rita
Williamson, Sinead
contents Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses the risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using vision-language models as a testbed, we present the first systematic comparison of information retained at different "representational levels'' as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks: low-dimensional projections of the residual stream obtained using tuned lens, and the final top-k logits most likely to impact model's answer. We show that even easily accessible bottlenecks defined by the model's top logit values can leak task-irrelevant information present in an image-based query, in some cases revealing as much information as direct projections of the full residual stream.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09885
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What do your logits know? (The answer may surprise you!)
Fedzechkina, Masha
Gualdoni, Eleonora
Ramos, Rita
Williamson, Sinead
Artificial Intelligence
Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses the risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using vision-language models as a testbed, we present the first systematic comparison of information retained at different "representational levels'' as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks: low-dimensional projections of the residual stream obtained using tuned lens, and the final top-k logits most likely to impact model's answer. We show that even easily accessible bottlenecks defined by the model's top logit values can leak task-irrelevant information present in an image-based query, in some cases revealing as much information as direct projections of the full residual stream.
title What do your logits know? (The answer may surprise you!)
topic Artificial Intelligence
url https://arxiv.org/abs/2604.09885