Does Transformer Interpretability Transfer to RNNs?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Paulo, Gonçalo, Marshall, Thomas, Belrose, Nora
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917634945056768
author Paulo, Gonçalo
Marshall, Thomas
Belrose, Nora
author_facet Paulo, Gonçalo
Marshall, Thomas
Belrose, Nora
contents Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of language modeling perplexity and downstream evaluations, suggesting that future systems may be built on completely new architectures. In this paper, we examine if selected interpretability methods originally designed for transformer language models will transfer to these up-and-coming recurrent architectures. Specifically, we focus on steering model outputs via contrastive activation addition, on eliciting latent predictions via the tuned lens, and eliciting latent knowledge from models fine-tuned to produce false outputs under certain conditions. Our results show that most of these techniques are effective when applied to RNNs, and we show that it is possible to improve some of them by taking advantage of RNNs' compressed state.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05971
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does Transformer Interpretability Transfer to RNNs?
Paulo, Gonçalo
Marshall, Thomas
Belrose, Nora
Machine Learning
Artificial Intelligence
Computation and Language
Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of language modeling perplexity and downstream evaluations, suggesting that future systems may be built on completely new architectures. In this paper, we examine if selected interpretability methods originally designed for transformer language models will transfer to these up-and-coming recurrent architectures. Specifically, we focus on steering model outputs via contrastive activation addition, on eliciting latent predictions via the tuned lens, and eliciting latent knowledge from models fine-tuned to produce false outputs under certain conditions. Our results show that most of these techniques are effective when applied to RNNs, and we show that it is possible to improve some of them by taking advantage of RNNs' compressed state.
title Does Transformer Interpretability Transfer to RNNs?
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2404.05971