How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Xinchi, Shen, William F., Chen, Yihong, Kurmanji, Meghdad, Cancedda, Nicola, Stenetorp, Pontus, Lane, Nicholas D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913729268940800
author Qiu, Xinchi
Shen, William F.
Chen, Yihong
Kurmanji, Meghdad
Cancedda, Nicola
Stenetorp, Pontus
Lane, Nicholas D.
author_facet Qiu, Xinchi
Shen, William F.
Chen, Yihong
Kurmanji, Meghdad
Cancedda, Nicola
Stenetorp, Pontus
Lane, Nicholas D.
contents While unlearning knowledge from large language models (LLMs) is receiving increasing attention, one important aspect remains unexplored. Existing approaches and benchmarks assume data points to-be-forgotten are independent, ignoring their inter-connectivity - a fundamental characteristic of real-world data structures. In this paper, we propose PISTOL, a method for compiling structural datasets. PISTOL leverages the inherently structured nature of contractual relationships, offering several key benefits. First, it enables insights into the impact of structural data on unlearning effectiveness. Second, it provides precise and concise ground truths for clearer evaluation. Third, its attribute generation does not require input from pre-trained LLMs, mitigating confounding risks. Leveraging datasets synthesized using PISTOL, we demonstrate how data inter-connectivity impacts LLM unlearning. Specifically, (a) in both the pre-trained and fine-tuned models, unlearning difficulty increases as data inter-connectivity grows, (b) there is a positive correlation between the density of the knowledge graph and unlearning difficulty, and (c) when the to-be-forgotten data is skewed towards one domain, balancing retaining performance across all domains is challenging.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16810
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
Qiu, Xinchi
Shen, William F.
Chen, Yihong
Kurmanji, Meghdad
Cancedda, Nicola
Stenetorp, Pontus
Lane, Nicholas D.
Machine Learning
Artificial Intelligence
Computation and Language
While unlearning knowledge from large language models (LLMs) is receiving increasing attention, one important aspect remains unexplored. Existing approaches and benchmarks assume data points to-be-forgotten are independent, ignoring their inter-connectivity - a fundamental characteristic of real-world data structures. In this paper, we propose PISTOL, a method for compiling structural datasets. PISTOL leverages the inherently structured nature of contractual relationships, offering several key benefits. First, it enables insights into the impact of structural data on unlearning effectiveness. Second, it provides precise and concise ground truths for clearer evaluation. Third, its attribute generation does not require input from pre-trained LLMs, mitigating confounding risks. Leveraging datasets synthesized using PISTOL, we demonstrate how data inter-connectivity impacts LLM unlearning. Specifically, (a) in both the pre-trained and fine-tuned models, unlearning difficulty increases as data inter-connectivity grows, (b) there is a positive correlation between the density of the knowledge graph and unlearning difficulty, and (c) when the to-be-forgotten data is skewed towards one domain, balancing retaining performance across all domains is challenging.
title How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.16810