From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Schiekiera, Louis, Zimmer, Max, Roux, Christophe, Pokutta, Sebastian, Günther, Fritz
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917275289780224
author Schiekiera, Louis
Zimmer, Max
Roux, Christophe
Pokutta, Sebastian
Günther, Fritz
author_facet Schiekiera, Louis
Zimmer, Max
Roux, Christophe
Pokutta, Sebastian
Günther, Fritz
contents We investigate the extent to which an LLM's hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer models, we run two experimental paradigms -- similarity-based forced choice and free association -- over a shared 5,000-word vocabulary, collecting 17.5M+ trials to build behavior-based similarity matrices. Using representational similarity analysis, we compare behavioral geometries to layerwise hidden-state similarity and benchmark against FastText, BERT, and cross-model consensus. We find that forced-choice behavior aligns substantially more with hidden-state geometry than free association. In a held-out-words regression, behavioral similarity (especially forced choice) predicts unseen hidden-state similarities beyond lexical baselines and cross-model consensus, indicating that behavior-only measurements retain recoverable information about internal semantic geometry. Finally, we discuss implications for the ability of behavioral tasks to uncover hidden cognitive states.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00628
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
Schiekiera, Louis
Zimmer, Max
Roux, Christophe
Pokutta, Sebastian
Günther, Fritz
Machine Learning
Artificial Intelligence
Computation and Language
We investigate the extent to which an LLM's hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer models, we run two experimental paradigms -- similarity-based forced choice and free association -- over a shared 5,000-word vocabulary, collecting 17.5M+ trials to build behavior-based similarity matrices. Using representational similarity analysis, we compare behavioral geometries to layerwise hidden-state similarity and benchmark against FastText, BERT, and cross-model consensus. We find that forced-choice behavior aligns substantially more with hidden-state geometry than free association. In a held-out-words regression, behavioral similarity (especially forced choice) predicts unseen hidden-state similarities beyond lexical baselines and cross-model consensus, indicating that behavior-only measurements retain recoverable information about internal semantic geometry. Finally, we discuss implications for the ability of behavioral tasks to uncover hidden cognitive states.
title From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.00628