OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Korznikov, Anton, Galichin, Andrey, Dontsov, Alexey, Rogov, Oleg, Tutubalina, Elena, Oseledets, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908560778067968
author Korznikov, Anton
Galichin, Andrey
Dontsov, Alexey
Rogov, Oleg
Tutubalina, Elena
Oseledets, Ivan
author_facet Korznikov, Anton
Galichin, Andrey
Dontsov, Alexey
Rogov, Oleg
Tutubalina, Elena
Oseledets, Ivan
contents Sparse autoencoders (SAEs) are a technique for sparse decomposition of neural network activations into human-interpretable features. However, current SAEs suffer from feature absorption, where specialized features capture instances of general features creating representation holes, and feature composition, where independent features merge into composite representations. In this work, we introduce Orthogonal SAE (OrtSAE), a novel approach aimed to mitigate these issues by enforcing orthogonality between the learned features. By implementing a new training procedure that penalizes high pairwise cosine similarity between SAE features, OrtSAE promotes the development of disentangled features while scaling linearly with the SAE size, avoiding significant computational overhead. We train OrtSAE across different models and layers and compare it with other methods. We find that OrtSAE discovers 9% more distinct features, reduces feature absorption (by 65%) and composition (by 15%), improves performance on spurious correlation removal (+6%), and achieves on-par performance for other downstream tasks compared to traditional SAEs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
Korznikov, Anton
Galichin, Andrey
Dontsov, Alexey
Rogov, Oleg
Tutubalina, Elena
Oseledets, Ivan
Machine Learning
Sparse autoencoders (SAEs) are a technique for sparse decomposition of neural network activations into human-interpretable features. However, current SAEs suffer from feature absorption, where specialized features capture instances of general features creating representation holes, and feature composition, where independent features merge into composite representations. In this work, we introduce Orthogonal SAE (OrtSAE), a novel approach aimed to mitigate these issues by enforcing orthogonality between the learned features. By implementing a new training procedure that penalizes high pairwise cosine similarity between SAE features, OrtSAE promotes the development of disentangled features while scaling linearly with the SAE size, avoiding significant computational overhead. We train OrtSAE across different models and layers and compare it with other methods. We find that OrtSAE discovers 9% more distinct features, reduces feature absorption (by 65%) and composition (by 15%), improves performance on spurious correlation removal (+6%), and achieves on-par performance for other downstream tasks compared to traditional SAEs.
title OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
topic Machine Learning
url https://arxiv.org/abs/2509.22033