Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yoon, Sangyeon, Hong, Hyesoo, Jeung, Wonje, No, Albert
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912872141946880
author Yoon, Sangyeon
Hong, Hyesoo
Jeung, Wonje
No, Albert
author_facet Yoon, Sangyeon
Hong, Hyesoo
Jeung, Wonje
No, Albert
contents Machine unlearning aims to remove specific content from trained models while preserving overall performance. However, the phenomenon of benign relearning, in which forgotten information reemerges even from benign fine-tuning data, reveals that existing unlearning methods remain fundamentally fragile. A common explanation attributes this effect to topical relevance, but we find this account insufficient. Through systematic analysis, we demonstrate that syntactic similarity, rather than topicality, is the primary driver: across benchmarks, syntactically similar data consistently trigger recovery even without topical overlap, due to their alignment in representations and gradients with the forgotten content. Motivated by this insight, we introduce syntactic diversification, which paraphrases the original forget queries into heterogeneous structures prior to unlearning. This approach effectively suppresses benign relearning, accelerates forgetting, and substantially alleviates the trade-off between unlearning efficacy and model utility.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03379
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
Yoon, Sangyeon
Hong, Hyesoo
Jeung, Wonje
No, Albert
Machine Learning
Artificial Intelligence
Machine unlearning aims to remove specific content from trained models while preserving overall performance. However, the phenomenon of benign relearning, in which forgotten information reemerges even from benign fine-tuning data, reveals that existing unlearning methods remain fundamentally fragile. A common explanation attributes this effect to topical relevance, but we find this account insufficient. Through systematic analysis, we demonstrate that syntactic similarity, rather than topicality, is the primary driver: across benchmarks, syntactically similar data consistently trigger recovery even without topical overlap, due to their alignment in representations and gradients with the forgotten content. Motivated by this insight, we introduce syntactic diversification, which paraphrases the original forget queries into heterogeneous structures prior to unlearning. This approach effectively suppresses benign relearning, accelerates forgetting, and substantially alleviates the trade-off between unlearning efficacy and model utility.
title Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03379