Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Csordás, Róbert, Potts, Christopher, Manning, Christopher D., Geiger, Atticus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911996503392256
author Csordás, Róbert
Potts, Christopher
Manning, Christopher D.
Geiger, Atticus
author_facet Csordás, Róbert
Potts, Christopher
Manning, Christopher D.
Geiger, Atticus
contents The Linear Representation Hypothesis (LRH) states that neural networks learn to encode concepts as directions in activation space, and a strong version of the LRH states that models learn only such encodings. In this paper, we present a counterexample to this strong LRH: when trained to repeat an input token sequence, gated recurrent neural networks (RNNs) learn to represent the token at each position with a particular order of magnitude, rather than a direction. These representations have layered features that are impossible to locate in distinct linear subspaces. To show this, we train interventions to predict and manipulate tokens by learning the scaling factor corresponding to each sequence position. These interventions indicate that the smallest RNNs find only this magnitude-based solution, while larger RNNs have linear representations. These findings strongly indicate that interpretability research should not be confined by the LRH.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10920
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
Csordás, Róbert
Potts, Christopher
Manning, Christopher D.
Geiger, Atticus
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
The Linear Representation Hypothesis (LRH) states that neural networks learn to encode concepts as directions in activation space, and a strong version of the LRH states that models learn only such encodings. In this paper, we present a counterexample to this strong LRH: when trained to repeat an input token sequence, gated recurrent neural networks (RNNs) learn to represent the token at each position with a particular order of magnitude, rather than a direction. These representations have layered features that are impossible to locate in distinct linear subspaces. To show this, we train interventions to predict and manipulate tokens by learning the scaling factor corresponding to each sequence position. These interventions indicate that the smallest RNNs find only this magnitude-based solution, while larger RNNs have linear representations. These findings strongly indicate that interpretability research should not be confined by the LRH.
title Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2408.10920