ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Yujie, Yang, Chengyi, Xiang, Zhishang, Song, Yiping, Su, Jinsong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914581335506944
author Lin, Yujie
Yang, Chengyi
Xiang, Zhishang
Song, Yiping
Su, Jinsong
author_facet Lin, Yujie
Yang, Chengyi
Xiang, Zhishang
Song, Yiping
Su, Jinsong
contents Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web corpora, raising concerns for privacy and safety. Existing machine unlearning methods primarily rely on retraining or aggressive fine-tuning, which are either computationally expensive or prone to degrading related knowledge and overall model utility. In this work, we reformulate machine unlearning as a precise knowledge re-mapping problem via model editing. We propose ZeroUnlearn, a few-shot unlearning framework. It overwrites sensitive inputs by mapping them to a neutral target state and removing their original representations. ZeroUnlearn enforces representational orthogonality through a multiplicative parameter update with a closed-form solution, enabling efficient and targeted unlearning. We further extend ZeroUnlearn to a gradient-based variant for multi-sample unlearning. Experiments demonstrate that our approach outperforms existing baselines while preserving general model utility. Our code is available at the github: https://github.com/XMUDeepLIT/ZeroUnlearn.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18879
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
Lin, Yujie
Yang, Chengyi
Xiang, Zhishang
Song, Yiping
Su, Jinsong
Machine Learning
Artificial Intelligence
Computation and Language
Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web corpora, raising concerns for privacy and safety. Existing machine unlearning methods primarily rely on retraining or aggressive fine-tuning, which are either computationally expensive or prone to degrading related knowledge and overall model utility. In this work, we reformulate machine unlearning as a precise knowledge re-mapping problem via model editing. We propose ZeroUnlearn, a few-shot unlearning framework. It overwrites sensitive inputs by mapping them to a neutral target state and removing their original representations. ZeroUnlearn enforces representational orthogonality through a multiplicative parameter update with a closed-form solution, enabling efficient and targeted unlearning. We further extend ZeroUnlearn to a gradient-based variant for multi-sample unlearning. Experiments demonstrate that our approach outperforms existing baselines while preserving general model utility. Our code is available at the github: https://github.com/XMUDeepLIT/ZeroUnlearn.
title ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.18879