Linear Representations of Hierarchical Concepts in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakata, Masaki, Heinzerling, Benjamin, Ito, Takumi, Yokoi, Sho, Inui, Kentaro
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908959538937856
author Sakata, Masaki
Heinzerling, Benjamin
Ito, Takumi
Yokoi, Sho
Inui, Kentaro
author_facet Sakata, Masaki
Heinzerling, Benjamin
Ito, Takumi
Yokoi, Sho
Inui, Kentaro
contents We investigate how and to what extent hierarchical relations (e.g., Japan $\subset$ Eastern Asia $\subset$ Asia) are encoded in the internal representations of language models. Building on Linear Relational Concepts, we train linear transformations specific to each hierarchical depth and semantic domain, and characterize representational differences associated with hierarchical relations by comparing these transformations. Going beyond prior work on the representational geometry of hierarchies in LMs, our analysis covers multi-token entities and cross-layer representations. Across multiple domains we learn such transformations and evaluate in-domain generalization to unseen data and cross-domain transfer. Experiments show that, within a domain, hierarchical relations can be linearly recovered from model representations. We then analyze how hierarchical information is encoded in representation space. We find that it is encoded in a relatively low-dimensional subspace and that this subspace tends to be domain-specific. Our main result is that hierarchy representation is highly similar across these domain-specific subspaces. Overall, we find that all models considered in our experiments encode concept hierarchies in the form of highly interpretable linear representations.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07886
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Linear Representations of Hierarchical Concepts in Language Models
Sakata, Masaki
Heinzerling, Benjamin
Ito, Takumi
Yokoi, Sho
Inui, Kentaro
Computation and Language
We investigate how and to what extent hierarchical relations (e.g., Japan $\subset$ Eastern Asia $\subset$ Asia) are encoded in the internal representations of language models. Building on Linear Relational Concepts, we train linear transformations specific to each hierarchical depth and semantic domain, and characterize representational differences associated with hierarchical relations by comparing these transformations. Going beyond prior work on the representational geometry of hierarchies in LMs, our analysis covers multi-token entities and cross-layer representations. Across multiple domains we learn such transformations and evaluate in-domain generalization to unseen data and cross-domain transfer. Experiments show that, within a domain, hierarchical relations can be linearly recovered from model representations. We then analyze how hierarchical information is encoded in representation space. We find that it is encoded in a relatively low-dimensional subspace and that this subspace tends to be domain-specific. Our main result is that hierarchy representation is highly similar across these domain-specific subspaces. Overall, we find that all models considered in our experiments encode concept hierarchies in the form of highly interpretable linear representations.
title Linear Representations of Hierarchical Concepts in Language Models
topic Computation and Language
url https://arxiv.org/abs/2604.07886