Saved in:
Bibliographic Details
Main Authors: Wang, Xu, Li, Zihao, Wang, Benyou, Hu, Yan, Zou, Difan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.24428
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908385982545920
author Wang, Xu
Li, Zihao
Wang, Benyou
Hu, Yan
Zou, Difan
author_facet Wang, Xu
Li, Zihao
Wang, Benyou
Hu, Yan
Zou, Difan
contents Large language models (LLMs) store vast amounts of information, making them powerful yet raising privacy and safety concerns when selective knowledge removal is required. Existing unlearning strategies, ranging from gradient-based fine-tuning and model editing to sparse autoencoder (SAE) steering, either lack interpretability or fail to provide a robust defense against adversarial prompts. We propose SAE-Guided Subspace Projection Unlearning (SSPU), a novel framework that leverages SAE features to drive targeted updates in the model's parameter space, enabling precise, interpretable, and robust unlearning. SSPU's three-stage pipeline performs data-driven layer and feature selection, subspace construction via QR decomposition, and constrained optimization that controls activations into an "irrelevant" subspace while preserving retained knowledge. Overall, we use SAE features to construct a subspace that supervises unlearning, refining the loss and adding a regularization term to guide interpretable parameter updates. In experiments on the WMDP-Cyber forget set and three utility benchmarks (MMLU, TruthfulQA, GSM8K), SSPU reduces harmful knowledge accuracy by 3.22% compared to the strongest baseline. It also improves adversarial robustness, lowering malicious accuracy under jailbreak prompts compared to baselines. Our findings expose the limitations of prior unlearning methods and demonstrate how interpretable subspace-guided optimization can achieve robust, controllable model behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Model Unlearning via Sparse Autoencoder Subspace Guided Projections
Wang, Xu
Li, Zihao
Wang, Benyou
Hu, Yan
Zou, Difan
Computation and Language
Machine Learning
Large language models (LLMs) store vast amounts of information, making them powerful yet raising privacy and safety concerns when selective knowledge removal is required. Existing unlearning strategies, ranging from gradient-based fine-tuning and model editing to sparse autoencoder (SAE) steering, either lack interpretability or fail to provide a robust defense against adversarial prompts. We propose SAE-Guided Subspace Projection Unlearning (SSPU), a novel framework that leverages SAE features to drive targeted updates in the model's parameter space, enabling precise, interpretable, and robust unlearning. SSPU's three-stage pipeline performs data-driven layer and feature selection, subspace construction via QR decomposition, and constrained optimization that controls activations into an "irrelevant" subspace while preserving retained knowledge. Overall, we use SAE features to construct a subspace that supervises unlearning, refining the loss and adding a regularization term to guide interpretable parameter updates. In experiments on the WMDP-Cyber forget set and three utility benchmarks (MMLU, TruthfulQA, GSM8K), SSPU reduces harmful knowledge accuracy by 3.22% compared to the strongest baseline. It also improves adversarial robustness, lowering malicious accuracy under jailbreak prompts compared to baselines. Our findings expose the limitations of prior unlearning methods and demonstrate how interpretable subspace-guided optimization can achieve robust, controllable model behavior.
title Model Unlearning via Sparse Autoencoder Subspace Guided Projections
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.24428