Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oldfield, James, Im, Shawn, Li, Sharon, Nicolaou, Mihalis A., Patras, Ioannis, Chrysos, Grigorios G
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909990074187776
author Oldfield, James
Im, Shawn
Li, Sharon
Nicolaou, Mihalis A.
Patras, Ioannis
Chrysos, Grigorios G
author_facet Oldfield, James
Im, Shawn
Li, Sharon
Nicolaou, Mihalis A.
Patras, Ioannis
Chrysos, Grigorios G
contents Multilayer perceptrons (MLPs) are an integral part of large language models, yet their dense representations render them difficult to understand, edit, and steer. Recent methods learn interpretable approximations via neuron-level sparsity, yet fail to faithfully reconstruct the original mapping--significantly increasing model's next-token cross-entropy loss. In this paper, we advocate for moving to layer-level sparsity to overcome the accuracy trade-off in sparse layer approximation. Under this paradigm, we introduce Mixture of Decoders (MxDs). MxDs generalize MLPs and Gated Linear Units, expanding pre-trained dense layers into tens of thousands of specialized sublayers. Through a flexible form of tensor factorization, each sparsely activating MxD sublayer implements a linear transformation with full-rank weights--preserving the original decoders' expressive capacity even under heavy sparsity. Experimentally, we show that MxDs significantly outperform state-of-the-art methods (e.g., Transcoders) on the sparsity-accuracy frontier in language models with up to 3B parameters. Further evaluations on sparse probing and feature steering demonstrate that MxDs learn similarly specialized features of natural language--opening up a promising new avenue for designing interpretable yet faithful decompositions. Our code is included at: https://github.com/james-oldfield/MxD/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21364
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
Oldfield, James
Im, Shawn
Li, Sharon
Nicolaou, Mihalis A.
Patras, Ioannis
Chrysos, Grigorios G
Machine Learning
Artificial Intelligence
Multilayer perceptrons (MLPs) are an integral part of large language models, yet their dense representations render them difficult to understand, edit, and steer. Recent methods learn interpretable approximations via neuron-level sparsity, yet fail to faithfully reconstruct the original mapping--significantly increasing model's next-token cross-entropy loss. In this paper, we advocate for moving to layer-level sparsity to overcome the accuracy trade-off in sparse layer approximation. Under this paradigm, we introduce Mixture of Decoders (MxDs). MxDs generalize MLPs and Gated Linear Units, expanding pre-trained dense layers into tens of thousands of specialized sublayers. Through a flexible form of tensor factorization, each sparsely activating MxD sublayer implements a linear transformation with full-rank weights--preserving the original decoders' expressive capacity even under heavy sparsity. Experimentally, we show that MxDs significantly outperform state-of-the-art methods (e.g., Transcoders) on the sparsity-accuracy frontier in language models with up to 3B parameters. Further evaluations on sparse probing and feature steering demonstrate that MxDs learn similarly specialized features of natural language--opening up a promising new avenue for designing interpretable yet faithful decompositions. Our code is included at: https://github.com/james-oldfield/MxD/.
title Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.21364