SIMU: Selective Influence Machine Unlearning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Agarwal, Anu, Pamnani, Mihir, Hakkani-Tur, Dilek
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908584193818624
author Agarwal, Anu
Pamnani, Mihir
Hakkani-Tur, Dilek
author_facet Agarwal, Anu
Pamnani, Mihir
Hakkani-Tur, Dilek
contents The undesired memorization of sensitive information by Large Language Models (LLMs) has emphasized the need for safety mechanisms that can regulate model behavior. This has led to the development of machine unlearning techniques that enable models to precisely forget sensitive and unwanted information. For machine unlearning, first-order and second-order optimizer-based methods have shown significant progress in enabling LLMs to forget targeted information. However, in doing so, these approaches often compromise the model's original capabilities, resulting in unlearned models that struggle to retain their prior knowledge and overall utility. To address this, we propose Selective Influence Machine Unlearning (SIMU), a two-step framework that enhances second-order optimizer-based unlearning by selectively updating only the critical neurons responsible for encoding the forget-set. By constraining updates to these targeted neurons, SIMU achieves comparable unlearning efficacy while substantially outperforming current methods in retaining the model's original knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SIMU: Selective Influence Machine Unlearning
Agarwal, Anu
Pamnani, Mihir
Hakkani-Tur, Dilek
Machine Learning
Artificial Intelligence
The undesired memorization of sensitive information by Large Language Models (LLMs) has emphasized the need for safety mechanisms that can regulate model behavior. This has led to the development of machine unlearning techniques that enable models to precisely forget sensitive and unwanted information. For machine unlearning, first-order and second-order optimizer-based methods have shown significant progress in enabling LLMs to forget targeted information. However, in doing so, these approaches often compromise the model's original capabilities, resulting in unlearned models that struggle to retain their prior knowledge and overall utility. To address this, we propose Selective Influence Machine Unlearning (SIMU), a two-step framework that enhances second-order optimizer-based unlearning by selectively updating only the critical neurons responsible for encoding the forget-set. By constraining updates to these targeted neurons, SIMU achieves comparable unlearning efficacy while substantially outperforming current methods in retaining the model's original knowledge.
title SIMU: Selective Influence Machine Unlearning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.07822