Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cho, Sungjun, Hwang, Dasol, Sala, Frederic, Hwang, Sangheum, Cho, Kyunghyun, Cha, Sungmin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908593745297408
author Cho, Sungjun
Hwang, Dasol
Sala, Frederic
Hwang, Sangheum
Cho, Kyunghyun
Cha, Sungmin
author_facet Cho, Sungjun
Hwang, Dasol
Sala, Frederic
Hwang, Sangheum
Cho, Kyunghyun
Cha, Sungmin
contents Current unlearning metrics for generative models evaluate success based on reference responses or classifier outputs rather than assessing the core objective: whether the unlearned model behaves indistinguishably from a model that never saw the unwanted data. This reference-specific approach creates systematic blind spots, allowing models to appear successful while retaining unwanted knowledge accessible through alternative prompts or attacks. We address these limitations by proposing Functional Alignment for Distributional Equivalence (FADE), a novel metric that measures distributional similarity between unlearned and reference models by comparing bidirectional likelihood assignments over generated samples. Unlike existing approaches that rely on predetermined references, FADE captures functional alignment across the entire output distribution, providing a principled assessment of genuine unlearning. Our experiments on the TOFU benchmark for LLM unlearning and the UnlearnCanvas benchmark for text-to-image diffusion model unlearning reveal that methods achieving near-optimal scores on traditional metrics fail to achieve distributional equivalence, with many becoming more distant from the gold standard than before unlearning. These findings expose fundamental gaps in current evaluation practices and demonstrate that FADE provides a more robust foundation for developing and assessing truly effective unlearning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12981
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
Cho, Sungjun
Hwang, Dasol
Sala, Frederic
Hwang, Sangheum
Cho, Kyunghyun
Cha, Sungmin
Machine Learning
Current unlearning metrics for generative models evaluate success based on reference responses or classifier outputs rather than assessing the core objective: whether the unlearned model behaves indistinguishably from a model that never saw the unwanted data. This reference-specific approach creates systematic blind spots, allowing models to appear successful while retaining unwanted knowledge accessible through alternative prompts or attacks. We address these limitations by proposing Functional Alignment for Distributional Equivalence (FADE), a novel metric that measures distributional similarity between unlearned and reference models by comparing bidirectional likelihood assignments over generated samples. Unlike existing approaches that rely on predetermined references, FADE captures functional alignment across the entire output distribution, providing a principled assessment of genuine unlearning. Our experiments on the TOFU benchmark for LLM unlearning and the UnlearnCanvas benchmark for text-to-image diffusion model unlearning reveal that methods achieving near-optimal scores on traditional metrics fail to achieve distributional equivalence, with many becoming more distant from the gold standard than before unlearning. These findings expose fundamental gaps in current evaluation practices and demonstrate that FADE provides a more robust foundation for developing and assessing truly effective unlearning methods.
title Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
topic Machine Learning
url https://arxiv.org/abs/2510.12981