Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Rongzhe, Niu, Peizhi, Hsu, Hans Hao-Hsun, Wu, Ruihan, Yin, Haoteng, Ghassemi, Mohsen, Li, Yifan, Potluru, Vamsi K., Chien, Eli, Chaudhuri, Kamalika, Milenkovic, Olgica, Li, Pan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911225042960384
author Wei, Rongzhe
Niu, Peizhi
Hsu, Hans Hao-Hsun
Wu, Ruihan
Yin, Haoteng
Ghassemi, Mohsen
Li, Yifan
Potluru, Vamsi K.
Chien, Eli
Chaudhuri, Kamalika
Milenkovic, Olgica
Li, Pan
author_facet Wei, Rongzhe
Niu, Peizhi
Hsu, Hans Hao-Hsun
Wu, Ruihan
Yin, Haoteng
Ghassemi, Mohsen
Li, Yifan
Potluru, Vamsi K.
Chien, Eli
Chaudhuri, Kamalika
Milenkovic, Olgica
Li, Pan
contents Machine unlearning techniques aim to mitigate unintended memorization in large language models (LLMs). However, existing approaches predominantly focus on the explicit removal of isolated facts, often overlooking latent inferential dependencies and the non-deterministic nature of knowledge within LLMs. Consequently, facts presumed forgotten may persist implicitly through correlated information. To address these challenges, we propose a knowledge unlearning evaluation framework that more accurately captures the implicit structure of real-world knowledge by representing relevant factual contexts as knowledge graphs with associated confidence scores. We further develop an inference-based evaluation protocol leveraging powerful LLMs as judges; these judges reason over the extracted knowledge subgraph to determine unlearning success. Our LLM judges utilize carefully designed prompts and are calibrated against human evaluations to ensure their trustworthiness and stability. Extensive experiments on our newly constructed benchmark demonstrate that our framework provides a more realistic and rigorous assessment of unlearning performance. Moreover, our findings reveal that current evaluation strategies tend to overestimate unlearning effectiveness. Our code is publicly available at https://github.com/Graph-COM/Knowledge_Unlearning.git.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
Wei, Rongzhe
Niu, Peizhi
Hsu, Hans Hao-Hsun
Wu, Ruihan
Yin, Haoteng
Ghassemi, Mohsen
Li, Yifan
Potluru, Vamsi K.
Chien, Eli
Chaudhuri, Kamalika
Milenkovic, Olgica
Li, Pan
Computation and Language
Machine Learning
Machine unlearning techniques aim to mitigate unintended memorization in large language models (LLMs). However, existing approaches predominantly focus on the explicit removal of isolated facts, often overlooking latent inferential dependencies and the non-deterministic nature of knowledge within LLMs. Consequently, facts presumed forgotten may persist implicitly through correlated information. To address these challenges, we propose a knowledge unlearning evaluation framework that more accurately captures the implicit structure of real-world knowledge by representing relevant factual contexts as knowledge graphs with associated confidence scores. We further develop an inference-based evaluation protocol leveraging powerful LLMs as judges; these judges reason over the extracted knowledge subgraph to determine unlearning success. Our LLM judges utilize carefully designed prompts and are calibrated against human evaluations to ensure their trustworthiness and stability. Extensive experiments on our newly constructed benchmark demonstrate that our framework provides a more realistic and rigorous assessment of unlearning performance. Moreover, our findings reveal that current evaluation strategies tend to overestimate unlearning effectiveness. Our code is publicly available at https://github.com/Graph-COM/Knowledge_Unlearning.git.
title Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.05735