Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Khoriaty, Matthew, Shportko, Andrii, Mercier, Gustavo, Wood-Doughty, Zach
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912274335137792
author Khoriaty, Matthew
Shportko, Andrii
Mercier, Gustavo
Wood-Doughty, Zach
author_facet Khoriaty, Matthew
Shportko, Andrii
Mercier, Gustavo
Wood-Doughty, Zach
contents Recent developments in Large Language Model (LLM) capabilities have brought great potential but also posed new risks. For example, LLMs with knowledge of bioweapons, advanced chemistry, or cyberattacks could cause violence if placed in the wrong hands or during malfunctions. Because of their nature as near-black boxes, intuitive interpretation of LLM internals remains an open research question, preventing developers from easily controlling model behavior and capabilities. The use of Sparse Autoencoders (SAEs) has recently emerged as a potential method of unraveling representations of concepts in LLMs internals, and has allowed developers to steer model outputs by directly modifying the hidden activations. In this paper, we use SAEs to identify unwanted concepts from the Weapons of Mass Destruction Proxy (WMDP) dataset within gemma-2-2b internals and use feature steering to reduce the model's ability to answer harmful questions while retaining its performance on harmless queries. Our results bring back optimism to the viability of SAE-based explicit knowledge unlearning techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
Khoriaty, Matthew
Shportko, Andrii
Mercier, Gustavo
Wood-Doughty, Zach
Machine Learning
Artificial Intelligence
Recent developments in Large Language Model (LLM) capabilities have brought great potential but also posed new risks. For example, LLMs with knowledge of bioweapons, advanced chemistry, or cyberattacks could cause violence if placed in the wrong hands or during malfunctions. Because of their nature as near-black boxes, intuitive interpretation of LLM internals remains an open research question, preventing developers from easily controlling model behavior and capabilities. The use of Sparse Autoencoders (SAEs) has recently emerged as a potential method of unraveling representations of concepts in LLMs internals, and has allowed developers to steer model outputs by directly modifying the hidden activations. In this paper, we use SAEs to identify unwanted concepts from the Weapons of Mass Destruction Proxy (WMDP) dataset within gemma-2-2b internals and use feature steering to reduce the model's ability to answer harmful questions while retaining its performance on harmless queries. Our results bring back optimism to the viability of SAE-based explicit knowledge unlearning techniques.
title Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.11127