Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Huazheng, Jing, Yongcheng, Sun, Haifeng, Wang, Yingjie, Wang, Jingyu, Liao, Jianxin, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909831877623808
author Wang, Huazheng
Jing, Yongcheng
Sun, Haifeng
Wang, Yingjie
Wang, Jingyu
Liao, Jianxin
Tao, Dacheng
author_facet Wang, Huazheng
Jing, Yongcheng
Sun, Haifeng
Wang, Yingjie
Wang, Jingyu
Liao, Jianxin
Tao, Dacheng
contents In this paper, we investigate knowledge forgetting in large language models with a focus on its generalisation, ensuring that models forget not only specific training samples but also related implicit knowledge. To this end, we begin by identifying a broader unlearning scope that includes both target data and logically associated samples, including rephrased, subject-replaced, relation-reversed, and one-hop reasoned data. We then conduct a rigorous evaluation of 15 state-of-the-art methods across three datasets, revealing that unlearned models still recall paraphrased answers and retain target facts in their intermediate layers. This motivates us to take a preliminary step toward more generalised implicit knowledge forgetting by proposing PerMU, a novel probability perturbation-based unlearning paradigm. PerMU simulates adversarial unlearning samples to eliminate fact-related tokens from the logit distribution, collectively reducing the probabilities of all answer-associated tokens. Experiments are conducted on a diverse range of datasets, including TOFU, Harry Potter, ZsRE, WMDP, and MUSE, using models ranging from 1.3B to 13B in scale. The results demonstrate that PerMU delivers up to a 50.40% improvement in unlearning vanilla target data while maintaining a 40.73% boost in forgetting implicit knowledge. Our code can be found in https://github.com/MaybeLizzy/PERMU.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19982
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
Wang, Huazheng
Jing, Yongcheng
Sun, Haifeng
Wang, Yingjie
Wang, Jingyu
Liao, Jianxin
Tao, Dacheng
Computation and Language
Machine Learning
In this paper, we investigate knowledge forgetting in large language models with a focus on its generalisation, ensuring that models forget not only specific training samples but also related implicit knowledge. To this end, we begin by identifying a broader unlearning scope that includes both target data and logically associated samples, including rephrased, subject-replaced, relation-reversed, and one-hop reasoned data. We then conduct a rigorous evaluation of 15 state-of-the-art methods across three datasets, revealing that unlearned models still recall paraphrased answers and retain target facts in their intermediate layers. This motivates us to take a preliminary step toward more generalised implicit knowledge forgetting by proposing PerMU, a novel probability perturbation-based unlearning paradigm. PerMU simulates adversarial unlearning samples to eliminate fact-related tokens from the logit distribution, collectively reducing the probabilities of all answer-associated tokens. Experiments are conducted on a diverse range of datasets, including TOFU, Harry Potter, ZsRE, WMDP, and MUSE, using models ranging from 1.3B to 13B in scale. The results demonstrate that PerMU delivers up to a 50.40% improvement in unlearning vanilla target data while maintaining a 40.73% boost in forgetting implicit knowledge. Our code can be found in https://github.com/MaybeLizzy/PERMU.
title Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.19982