Understanding Gated Neurons in Transformers from Their Input-Output Functionality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gerstner, Sebastian, Schütze, Hinrich
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916754836422656
author Gerstner, Sebastian
Schütze, Hinrich
author_facet Gerstner, Sebastian
Schütze, Hinrich
contents Interpretability researchers have attempted to understand MLP neurons of language models based on both the contexts in which they activate and their output weight vectors. They have paid little attention to a complementary aspect: the interactions between input and output. For example, when neurons detect a direction in the input, they might add much the same direction to the residual stream ("enrichment neurons") or reduce its presence ("depletion neurons"). We address this aspect by examining the cosine similarity between input and output weights of a neuron. We apply our method to 12 models and find that enrichment neurons dominate in early-middle layers whereas later layers tend more towards depletion. To explain this finding, we argue that enrichment neurons are largely responsible for enriching concept representations, one of the first steps of factual recall. Our input-output perspective is a complement to activation-dependent analyses and to approaches that treat input and output separately.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Gated Neurons in Transformers from Their Input-Output Functionality
Gerstner, Sebastian
Schütze, Hinrich
Machine Learning
Computation and Language
Interpretability researchers have attempted to understand MLP neurons of language models based on both the contexts in which they activate and their output weight vectors. They have paid little attention to a complementary aspect: the interactions between input and output. For example, when neurons detect a direction in the input, they might add much the same direction to the residual stream ("enrichment neurons") or reduce its presence ("depletion neurons"). We address this aspect by examining the cosine similarity between input and output weights of a neuron. We apply our method to 12 models and find that enrichment neurons dominate in early-middle layers whereas later layers tend more towards depletion. To explain this finding, we argue that enrichment neurons are largely responsible for enriching concept representations, one of the first steps of factual recall. Our input-output perspective is a complement to activation-dependent analyses and to approaches that treat input and output separately.
title Understanding Gated Neurons in Transformers from Their Input-Output Functionality
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.17936