Semantic Structure in Large Language Model Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kozlowski, Austin C., Dai, Callin, Boutyline, Andrei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908489426665472
author Kozlowski, Austin C.
Dai, Callin
Boutyline, Andrei
author_facet Kozlowski, Austin C.
Dai, Callin
Boutyline, Andrei
contents Psychological research consistently finds that human ratings of words across diverse semantic scales can be reduced to a low-dimensional form with relatively little information loss. We find that the semantic associations encoded in the embedding matrices of large language models (LLMs) exhibit a similar structure. We show that the projections of words on semantic directions defined by antonym pairs (e.g. kind - cruel) correlate highly with human ratings, and further find that these projections effectively reduce to a 3-dimensional subspace within LLM embeddings, closely resembling the patterns derived from human survey responses. Moreover, we find that shifting tokens along one semantic direction causes off-target effects on geometrically aligned features proportional to their cosine similarity. These findings suggest that semantic features are entangled within LLMs similarly to how they are interconnected in human language, and a great deal of semantic information, despite its apparent complexity, is surprisingly low-dimensional. Furthermore, accounting for this semantic structure may prove essential for avoiding unintended consequences when steering features.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10003
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic Structure in Large Language Model Embeddings
Kozlowski, Austin C.
Dai, Callin
Boutyline, Andrei
Computation and Language
Artificial Intelligence
Psychological research consistently finds that human ratings of words across diverse semantic scales can be reduced to a low-dimensional form with relatively little information loss. We find that the semantic associations encoded in the embedding matrices of large language models (LLMs) exhibit a similar structure. We show that the projections of words on semantic directions defined by antonym pairs (e.g. kind - cruel) correlate highly with human ratings, and further find that these projections effectively reduce to a 3-dimensional subspace within LLM embeddings, closely resembling the patterns derived from human survey responses. Moreover, we find that shifting tokens along one semantic direction causes off-target effects on geometrically aligned features proportional to their cosine similarity. These findings suggest that semantic features are entangled within LLMs similarly to how they are interconnected in human language, and a great deal of semantic information, despite its apparent complexity, is surprisingly low-dimensional. Furthermore, accounting for this semantic structure may prove essential for avoiding unintended consequences when steering features.
title Semantic Structure in Large Language Model Embeddings
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.10003