Reveal and Release: Iterative LLM Unlearning with Self-generated Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xie, Linxi, Teng, Xin, Ke, Shichang, Wen, Hongyi, Wang, Shengjie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915500907298816
author Xie, Linxi
Teng, Xin
Ke, Shichang
Wen, Hongyi
Wang, Shengjie
author_facet Xie, Linxi
Teng, Xin
Ke, Shichang
Wen, Hongyi
Wang, Shengjie
contents Large language model (LLM) unlearning has demonstrated effectiveness in removing the influence of undesirable data (also known as forget data). Existing approaches typically assume full access to the forget dataset, overlooking two key challenges: (1) Forget data is often privacy-sensitive, rare, or legally regulated, making it expensive or impractical to obtain (2) The distribution of available forget data may not align with how that information is represented within the model. To address these limitations, we propose a ``Reveal-and-Release'' method to unlearn with self-generated data, where we prompt the model to reveal what it knows using optimized instructions. To fully utilize the self-generated forget data, we propose an iterative unlearning framework, where we make incremental adjustments to the model's weight space with parameter-efficient modules trained on the forget data. Experimental results demonstrate that our method balances the tradeoff between forget quality and utility preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14624
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reveal and Release: Iterative LLM Unlearning with Self-generated Data
Xie, Linxi
Teng, Xin
Ke, Shichang
Wen, Hongyi
Wang, Shengjie
Computation and Language
Artificial Intelligence
Machine Learning
Large language model (LLM) unlearning has demonstrated effectiveness in removing the influence of undesirable data (also known as forget data). Existing approaches typically assume full access to the forget dataset, overlooking two key challenges: (1) Forget data is often privacy-sensitive, rare, or legally regulated, making it expensive or impractical to obtain (2) The distribution of available forget data may not align with how that information is represented within the model. To address these limitations, we propose a ``Reveal-and-Release'' method to unlearn with self-generated data, where we prompt the model to reveal what it knows using optimized instructions. To fully utilize the self-generated forget data, we propose an iterative unlearning framework, where we make incremental adjustments to the model's weight space with parameter-efficient modules trained on the forget data. Experimental results demonstrate that our method balances the tradeoff between forget quality and utility preservation.
title Reveal and Release: Iterative LLM Unlearning with Self-generated Data
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.14624