Logit Distance Bounds Representational Similarity

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nielsen, Beatrix M. G., Marconato, Emanuele, Gresele, Luigi, Dittadi, Andrea, Buchholz, Simon
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914338150809600
author Nielsen, Beatrix M. G.
Marconato, Emanuele
Gresele, Luigi
Dittadi, Andrea
Buchholz, Simon
author_facet Nielsen, Beatrix M. G.
Marconato, Emanuele
Gresele, Luigi
Dittadi, Andrea
Buchholz, Simon
contents For a broad family of discriminative models that includes autoregressive language models, identifiability results imply that if two models induce the same conditional distributions, then their internal representations agree up to an invertible linear transformation. We ask whether an analogous conclusion holds approximately when the distributions are close instead of equal. Building on the observation of Nielsen et al. (2025) that closeness in KL divergence need not imply high linear representational similarity, we study a distributional distance based on logit differences and show that closeness in this distance does yield linear similarity guarantees. Specifically, we define a representational dissimilarity measure based on the models' identifiability class and prove that it is bounded by the logit distance. We further show that, when model probabilities are bounded away from zero, KL divergence upper-bounds logit distance; yet the resulting bound fails to provide nontrivial control in practice. As a consequence, KL-based distillation can match a teacher's predictions while failing to preserve linear representational properties, such as linear-probe recoverability of human-interpretable concepts. In distillation experiments on synthetic and image datasets, logit-distance distillation yields students with higher linear representational similarity and better preservation of the teacher's linearly recoverable concepts.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15438
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Logit Distance Bounds Representational Similarity
Nielsen, Beatrix M. G.
Marconato, Emanuele
Gresele, Luigi
Dittadi, Andrea
Buchholz, Simon
Machine Learning
Artificial Intelligence
For a broad family of discriminative models that includes autoregressive language models, identifiability results imply that if two models induce the same conditional distributions, then their internal representations agree up to an invertible linear transformation. We ask whether an analogous conclusion holds approximately when the distributions are close instead of equal. Building on the observation of Nielsen et al. (2025) that closeness in KL divergence need not imply high linear representational similarity, we study a distributional distance based on logit differences and show that closeness in this distance does yield linear similarity guarantees. Specifically, we define a representational dissimilarity measure based on the models' identifiability class and prove that it is bounded by the logit distance. We further show that, when model probabilities are bounded away from zero, KL divergence upper-bounds logit distance; yet the resulting bound fails to provide nontrivial control in practice. As a consequence, KL-based distillation can match a teacher's predictions while failing to preserve linear representational properties, such as linear-probe recoverability of human-interpretable concepts. In distillation experiments on synthetic and image datasets, logit-distance distillation yields students with higher linear representational similarity and better preservation of the teacher's linearly recoverable concepts.
title Logit Distance Bounds Representational Similarity
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.15438