Applying sparse autoencoders to unlearn knowledge in language models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Farrell, Eoin, Lau, Yeu-Tong, Conmy, Arthur
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916465620287488
author Farrell, Eoin
Lau, Yeu-Tong
Conmy, Arthur
author_facet Farrell, Eoin
Lau, Yeu-Tong
Conmy, Arthur
contents We investigate whether sparse autoencoders (SAEs) can be used to remove knowledge from language models. We use the biology subset of the Weapons of Mass Destruction Proxy dataset and test on the gemma-2b-it and gemma-2-2b-it language models. We demonstrate that individual interpretable biology-related SAE features can be used to unlearn a subset of WMDP-Bio questions with minimal side-effects in domains other than biology. Our results suggest that negative scaling of feature activations is necessary and that zero ablating features is ineffective. We find that intervening using multiple SAE features simultaneously can unlearn multiple different topics, but with similar or larger unwanted side-effects than the existing Representation Misdirection for Unlearning technique. Current SAE quality or intervention techniques would need to improve to make SAE-based unlearning comparable to the existing fine-tuning based techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19278
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Applying sparse autoencoders to unlearn knowledge in language models
Farrell, Eoin
Lau, Yeu-Tong
Conmy, Arthur
Machine Learning
Artificial Intelligence
We investigate whether sparse autoencoders (SAEs) can be used to remove knowledge from language models. We use the biology subset of the Weapons of Mass Destruction Proxy dataset and test on the gemma-2b-it and gemma-2-2b-it language models. We demonstrate that individual interpretable biology-related SAE features can be used to unlearn a subset of WMDP-Bio questions with minimal side-effects in domains other than biology. Our results suggest that negative scaling of feature activations is necessary and that zero ablating features is ineffective. We find that intervening using multiple SAE features simultaneously can unlearn multiple different topics, but with similar or larger unwanted side-effects than the existing Representation Misdirection for Unlearning technique. Current SAE quality or intervention techniques would need to improve to make SAE-based unlearning comparable to the existing fine-tuning based techniques.
title Applying sparse autoencoders to unlearn knowledge in language models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.19278