Polar probe linearly decodes semantic structures from LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Diego-Simón, Pablo J., Orhan, Pierre, Chemla, Emmanuel, Lakretz, Yair, King, Jean-Rémi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911690180788224
author Diego-Simón, Pablo J.
Orhan, Pierre
Chemla, Emmanuel
Lakretz, Yair
King, Jean-Rémi
author_facet Diego-Simón, Pablo J.
Orhan, Pierre
Chemla, Emmanuel
Lakretz, Yair
King, Jean-Rémi
contents How do artificial neural networks bind concepts to form complex semantic structures? Here, we propose a simple neural code, whereby the existence and the type of relations between entities are represented by the distance and the direction between their embeddings, respectively. We test this hypothesis in a variety of Large Language Models (LLMs), each input with natural-language descriptions of minimalist tasks from five different domains: arithmetic, visual scenes, family trees, metro maps and social interactions. Results show that the true semantic structures can be linearly recovered with a Polar Probe targeting a subspace of LLMs' layer activations. Second, this code emerges mostly in middle layers and improves with LLM performance. Third, these Polar Probes successfully generalize to new entities and relation types, but degrades with the size of the semantic structure. Finally, the quality of the polar representation correlates with the LLM's ability to answer questions about the semantic structure. Together, these findings suggest that LLMs learn to build complex semantic structures by binding representations with a simple geometrical principle.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14125
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Polar probe linearly decodes semantic structures from LLMs
Diego-Simón, Pablo J.
Orhan, Pierre
Chemla, Emmanuel
Lakretz, Yair
King, Jean-Rémi
Computation and Language
How do artificial neural networks bind concepts to form complex semantic structures? Here, we propose a simple neural code, whereby the existence and the type of relations between entities are represented by the distance and the direction between their embeddings, respectively. We test this hypothesis in a variety of Large Language Models (LLMs), each input with natural-language descriptions of minimalist tasks from five different domains: arithmetic, visual scenes, family trees, metro maps and social interactions. Results show that the true semantic structures can be linearly recovered with a Polar Probe targeting a subspace of LLMs' layer activations. Second, this code emerges mostly in middle layers and improves with LLM performance. Third, these Polar Probes successfully generalize to new entities and relation types, but degrades with the size of the semantic structure. Finally, the quality of the polar representation correlates with the LLM's ability to answer questions about the semantic structure. Together, these findings suggest that LLMs learn to build complex semantic structures by binding representations with a simple geometrical principle.
title Polar probe linearly decodes semantic structures from LLMs
topic Computation and Language
url https://arxiv.org/abs/2605.14125