Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Karvonen, Adam, Wright, Benjamin, Rager, Can, Angell, Rico, Brinkmann, Jannik, Smith, Logan, Verdun, Claudio Mayrink, Bau, David, Marks, Samuel
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909371355627520
author Karvonen, Adam
Wright, Benjamin
Rager, Can
Angell, Rico
Brinkmann, Jannik
Smith, Logan
Verdun, Claudio Mayrink
Bau, David
Marks, Samuel
author_facet Karvonen, Adam
Wright, Benjamin
Rager, Can
Angell, Rico
Brinkmann, Jannik
Smith, Logan
Verdun, Claudio Mayrink
Bau, David
Marks, Samuel
contents What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features that we expect good SAEs to recover. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on chess and Othello transcripts. These settings carry natural collections of interpretable features -- for example, "there is a knight on F3" -- which we leverage into $\textit{supervised}$ metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $\textit{p-annealing}$, which improves performance on prior unsupervised metrics as well as our new metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
Karvonen, Adam
Wright, Benjamin
Rager, Can
Angell, Rico
Brinkmann, Jannik
Smith, Logan
Verdun, Claudio Mayrink
Bau, David
Marks, Samuel
Machine Learning
Artificial Intelligence
Computation and Language
What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features that we expect good SAEs to recover. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on chess and Othello transcripts. These settings carry natural collections of interpretable features -- for example, "there is a knight on F3" -- which we leverage into $\textit{supervised}$ metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $\textit{p-annealing}$, which improves performance on prior unsupervised metrics as well as our new metrics.
title Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2408.00113