Large Language Models Encode Semantics and Alignment in Linearly Separable Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saglam, Baturay, Kassianik, Paul, Nelson, Blaine, Weerawardhena, Sajana, Singer, Yaron, Karbasi, Amin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908777841688576
author Saglam, Baturay
Kassianik, Paul
Nelson, Blaine
Weerawardhena, Sajana
Singer, Yaron
Karbasi, Amin
author_facet Saglam, Baturay
Kassianik, Paul
Nelson, Blaine
Weerawardhena, Sajana
Singer, Yaron
Karbasi, Amin
contents Understanding the latent space geometry of large language models (LLMs) is key to interpreting their behavior and improving alignment. Yet it remains unclear to what extent LLMs linearly organize representations related to semantic understanding. To explore this, we conduct a large-scale empirical study of hidden representations in 11 autoregressive models across six scientific topics. We find that high-level semantic information consistently resides in low-dimensional subspaces that form linearly separable representations across domains. This separability becomes more pronounced in deeper layers and under prompts that elicit structured reasoning or alignment behavior$\unicode{x2013}$even when surface content remains unchanged. These findings motivate geometry-aware tools that operate directly in latent space to detect and mitigate harmful and adversarial content. As a proof of concept, we train an MLP probe on final-layer hidden states as a lightweight latent-space guardrail. This approach substantially improves refusal rates on malicious queries and prompt injections that bypass both the model's built-in safety alignment and external token-level filters.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09709
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
Saglam, Baturay
Kassianik, Paul
Nelson, Blaine
Weerawardhena, Sajana
Singer, Yaron
Karbasi, Amin
Computation and Language
Machine Learning
Understanding the latent space geometry of large language models (LLMs) is key to interpreting their behavior and improving alignment. Yet it remains unclear to what extent LLMs linearly organize representations related to semantic understanding. To explore this, we conduct a large-scale empirical study of hidden representations in 11 autoregressive models across six scientific topics. We find that high-level semantic information consistently resides in low-dimensional subspaces that form linearly separable representations across domains. This separability becomes more pronounced in deeper layers and under prompts that elicit structured reasoning or alignment behavior$\unicode{x2013}$even when surface content remains unchanged. These findings motivate geometry-aware tools that operate directly in latent space to detect and mitigate harmful and adversarial content. As a proof of concept, we train an MLP probe on final-layer hidden states as a lightweight latent-space guardrail. This approach substantially improves refusal rates on malicious queries and prompt injections that bypass both the model's built-in safety alignment and external token-level filters.
title Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.09709