The Origins of Representation Manifolds in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Modell, Alexander, Rubin-Delanchy, Patrick, Whiteley, Nick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916756318060544
author Modell, Alexander
Rubin-Delanchy, Patrick
Whiteley, Nick
author_facet Modell, Alexander
Rubin-Delanchy, Patrick
Whiteley, Nick
contents There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key element of this effort is the linear representation hypothesis, which posits that neural representations are sparse linear combinations of `almost-orthogonal' direction vectors, reflecting the presence or absence of different features. This model underpins the use of sparse autoencoders to recover features from representations. Moving towards a fuller model of features, in which neural representations could encode not just the presence but also a potentially continuous and multidimensional value for a feature, has been a subject of intense recent discourse. We describe why and how a feature might be represented as a manifold, demonstrating in particular that cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths, potentially answering the question of how distance in representation space and relatedness in concept space could be connected. The critical assumptions and predictions of the theory are validated on text embeddings and token activations of large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Origins of Representation Manifolds in Large Language Models
Modell, Alexander
Rubin-Delanchy, Patrick
Whiteley, Nick
Machine Learning
Artificial Intelligence
68T07
I.2.7
There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key element of this effort is the linear representation hypothesis, which posits that neural representations are sparse linear combinations of `almost-orthogonal' direction vectors, reflecting the presence or absence of different features. This model underpins the use of sparse autoencoders to recover features from representations. Moving towards a fuller model of features, in which neural representations could encode not just the presence but also a potentially continuous and multidimensional value for a feature, has been a subject of intense recent discourse. We describe why and how a feature might be represented as a manifold, demonstrating in particular that cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths, potentially answering the question of how distance in representation space and relatedness in concept space could be connected. The critical assumptions and predictions of the theory are validated on text embeddings and token activations of large language models.
title The Origins of Representation Manifolds in Large Language Models
topic Machine Learning
Artificial Intelligence
68T07
I.2.7
url https://arxiv.org/abs/2505.18235