Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909371355627520 |
|---|---|
| author | Karvonen, Adam Wright, Benjamin Rager, Can Angell, Rico Brinkmann, Jannik Smith, Logan Verdun, Claudio Mayrink Bau, David Marks, Samuel |
| author_facet | Karvonen, Adam Wright, Benjamin Rager, Can Angell, Rico Brinkmann, Jannik Smith, Logan Verdun, Claudio Mayrink Bau, David Marks, Samuel |
| contents | What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features that we expect good SAEs to recover. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on chess and Othello transcripts. These settings carry natural collections of interpretable features -- for example, "there is a knight on F3" -- which we leverage into $\textit{supervised}$ metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $\textit{p-annealing}$, which improves performance on prior unsupervised metrics as well as our new metrics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_00113 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models Karvonen, Adam Wright, Benjamin Rager, Can Angell, Rico Brinkmann, Jannik Smith, Logan Verdun, Claudio Mayrink Bau, David Marks, Samuel Machine Learning Artificial Intelligence Computation and Language What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features that we expect good SAEs to recover. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on chess and Othello transcripts. These settings carry natural collections of interpretable features -- for example, "there is a knight on F3" -- which we leverage into $\textit{supervised}$ metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $\textit{p-annealing}$, which improves performance on prior unsupervised metrics as well as our new metrics. |
| title | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2408.00113 |