Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muchane, Mark, Richardson, Sean, Park, Kiho, Veitch, Victor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910980062052352
author Muchane, Mark
Richardson, Sean
Park, Kiho
Veitch, Victor
author_facet Muchane, Mark
Richardson, Sean
Park, Kiho
Veitch, Victor
contents Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01197
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
Muchane, Mark
Richardson, Sean
Park, Kiho
Veitch, Victor
Computation and Language
Artificial Intelligence
Machine Learning
Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.
title Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.01197