Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Lingzhi, Zeng, Xingshan, Guo, Jinsong, Wong, Kam-Fai, Gottlob, Georg
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916524008144896
author Wang, Lingzhi
Zeng, Xingshan
Guo, Jinsong
Wong, Kam-Fai
Gottlob, Georg
author_facet Wang, Lingzhi
Zeng, Xingshan
Guo, Jinsong
Wong, Kam-Fai
Gottlob, Georg
contents This paper explores Machine Unlearning (MU), an emerging field that is gaining increased attention due to concerns about neural models unintentionally remembering personal or sensitive information. We present SeUL, a novel method that enables selective and fine-grained unlearning for language models. Unlike previous work that employs a fully reversed training objective in unlearning, SeUL minimizes the negative impact on the capability of language models, particularly in terms of generation. Furthermore, we introduce two innovative evaluation metrics, sensitive extraction likelihood (S-EL) and sensitive memorization accuracy (S-MA), specifically designed to assess the effectiveness of forgetting sensitive information. In support of the unlearning framework, we propose efficient automatic online and offline sensitive span annotation methods. The online selection method, based on language probability scores, ensures computational efficiency, while the offline annotation involves a two-stage LLM-based process for robust verification. In summary, this paper contributes a novel selective unlearning method (SeUL), introduces specialized evaluation metrics (S-EL and S-MA) for assessing sensitive information forgetting, and proposes automatic online and offline sensitive span annotation methods to support the overall unlearning framework and evaluation process.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models
Wang, Lingzhi
Zeng, Xingshan
Guo, Jinsong
Wong, Kam-Fai
Gottlob, Georg
Computation and Language
Artificial Intelligence
This paper explores Machine Unlearning (MU), an emerging field that is gaining increased attention due to concerns about neural models unintentionally remembering personal or sensitive information. We present SeUL, a novel method that enables selective and fine-grained unlearning for language models. Unlike previous work that employs a fully reversed training objective in unlearning, SeUL minimizes the negative impact on the capability of language models, particularly in terms of generation. Furthermore, we introduce two innovative evaluation metrics, sensitive extraction likelihood (S-EL) and sensitive memorization accuracy (S-MA), specifically designed to assess the effectiveness of forgetting sensitive information. In support of the unlearning framework, we propose efficient automatic online and offline sensitive span annotation methods. The online selection method, based on language probability scores, ensures computational efficiency, while the offline annotation involves a two-stage LLM-based process for robust verification. In summary, this paper contributes a novel selective unlearning method (SeUL), introduces specialized evaluation metrics (S-EL and S-MA) for assessing sensitive information forgetting, and proposes automatic online and offline sensitive span annotation methods to support the overall unlearning framework and evaluation process.
title Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.05813