GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Kun-Woo, Park, Ji-Hoon, Han, Ju-Min, Lee, Seong-Whan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909582322827264
author Kim, Kun-Woo
Park, Ji-Hoon
Han, Ju-Min
Lee, Seong-Whan
author_facet Kim, Kun-Woo
Park, Ji-Hoon
Han, Ju-Min
Lee, Seong-Whan
contents Large Language Models (LLMs) trained on extensive datasets often learn sensitive information, which raises significant social and legal concerns under principles such as the "Right to be forgotten." Retraining entire models from scratch to remove undesired information is both costly and impractical. Furthermore, existing single-domain unlearning methods fail to address multi-domain scenarios, where knowledge is interwoven across domains such as privacy and copyright, creating overlapping representations that lead to excessive knowledge removal or degraded performance. To tackle these issues, we propose GRAIL (GRadient-based AdaptIve unLearning), a novel multi-domain unlearning framework. GRAIL leverages gradient information from multiple domains to precisely distinguish the unlearning scope from the retention scope, and applies an adaptive parameter-wise localization strategy to selectively remove targeted knowledge while preserving critical parameters for each domain. Experimental results on unlearning benchmarks show that GRAIL achieves unlearning success on par with the existing approaches, while also demonstrating up to 17% stronger knowledge retention success compared to the previous state-of-art method. Our findings establish a new paradigm for effectively managing and regulating sensitive information in large-scale pre-trained language models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
Kim, Kun-Woo
Park, Ji-Hoon
Han, Ju-Min
Lee, Seong-Whan
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) trained on extensive datasets often learn sensitive information, which raises significant social and legal concerns under principles such as the "Right to be forgotten." Retraining entire models from scratch to remove undesired information is both costly and impractical. Furthermore, existing single-domain unlearning methods fail to address multi-domain scenarios, where knowledge is interwoven across domains such as privacy and copyright, creating overlapping representations that lead to excessive knowledge removal or degraded performance. To tackle these issues, we propose GRAIL (GRadient-based AdaptIve unLearning), a novel multi-domain unlearning framework. GRAIL leverages gradient information from multiple domains to precisely distinguish the unlearning scope from the retention scope, and applies an adaptive parameter-wise localization strategy to selectively remove targeted knowledge while preserving critical parameters for each domain. Experimental results on unlearning benchmarks show that GRAIL achieves unlearning success on par with the existing approaches, while also demonstrating up to 17% stronger knowledge retention success compared to the previous state-of-art method. Our findings establish a new paradigm for effectively managing and regulating sensitive information in large-scale pre-trained language models.
title GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.12681