Binary Autoencoder for Mechanistic Interpretability of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Hakaze, Yang, Haolin, Li, Yanshu, Kurkoski, Brian M., Inoue, Naoya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915793535500288
author Cho, Hakaze
Yang, Haolin
Li, Yanshu
Kurkoski, Brian M.
Inoue, Naoya
author_facet Cho, Hakaze
Yang, Haolin
Li, Yanshu
Kurkoski, Brian M.
Inoue, Naoya
contents Existing works are dedicated to untangling atomized numerical components (features) from the hidden states of Large Language Models (LLMs). However, they typically rely on autoencoders constrained by some training-time regularization on single training instances, without an explicit guarantee of global sparsity among instances, causing a large amount of dense (simultaneously inactive) features, harming the feature sparsity and atomization. In this paper, we propose a novel autoencoder variant that enforces minimal entropy on minibatches of hidden activations, thereby promoting feature independence and sparsity across instances. For efficient entropy calculation, we discretize the hidden activations to 1-bit via a step function and apply gradient estimation to enable backpropagation, so that we term it as Binary Autoencoder (BAE) and empirically demonstrate two major applications: (1) Feature set entropy calculation. Entropy can be reliably estimated on binary hidden activations, which can be leveraged to characterize the inference dynamics of LLMs. (2) Feature untangling. Compared to typical methods, due to improved training strategy, BAE avoids dense features while producing the largest number of interpretable ones among baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20997
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Binary Autoencoder for Mechanistic Interpretability of Large Language Models
Cho, Hakaze
Yang, Haolin
Li, Yanshu
Kurkoski, Brian M.
Inoue, Naoya
Machine Learning
Artificial Intelligence
Computation and Language
Existing works are dedicated to untangling atomized numerical components (features) from the hidden states of Large Language Models (LLMs). However, they typically rely on autoencoders constrained by some training-time regularization on single training instances, without an explicit guarantee of global sparsity among instances, causing a large amount of dense (simultaneously inactive) features, harming the feature sparsity and atomization. In this paper, we propose a novel autoencoder variant that enforces minimal entropy on minibatches of hidden activations, thereby promoting feature independence and sparsity across instances. For efficient entropy calculation, we discretize the hidden activations to 1-bit via a step function and apply gradient estimation to enable backpropagation, so that we term it as Binary Autoencoder (BAE) and empirically demonstrate two major applications: (1) Feature set entropy calculation. Entropy can be reliably estimated on binary hidden activations, which can be leveraged to characterize the inference dynamics of LLMs. (2) Feature untangling. Compared to typical methods, due to improved training strategy, BAE avoids dense features while producing the largest number of interpretable ones among baselines.
title Binary Autoencoder for Mechanistic Interpretability of Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.20997