A Unified Framework for Diffusion Model Unlearning with f-Divergence

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Novello, Nicola, Fontana, Federico, Cinque, Luigi, Gunduz, Deniz, Tonello, Andrea M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913163136466944
author Novello, Nicola
Fontana, Federico
Cinque, Luigi
Gunduz, Deniz
Tonello, Andrea M.
author_facet Novello, Nicola
Fontana, Federico
Cinque, Luigi
Gunduz, Deniz
Tonello, Andrea M.
contents Most existing methods for concept unlearning in text-to-image diffusion models minimize a mean squared error (MSE) loss between the denoiser outputs conditioned on a target and an anchor concept, which is implicitly the KL divergence between two Gaussians. We generalize this objective to any $f$-divergence, recovering MSE as the KL instance, and identify a family of $α$-divergences whose Gaussian closed-form yields cheap, MSE-like training objectives. For the remaining $f$-divergences, we provide a min-max objective based on the variational formulation of the $f$-divergence. We theoretically analyze and numerically validate how different $f$-divergences impact the gradient magnitude and the convergence properties of the algorithm, affecting the quality of unlearning. For instance, we observe that the Hellinger closed-form instance consistently dominates MSE across multiple scenarios. More generally, the proposed unified framework offers a flexible paradigm for selecting the optimal divergence based on the application and user goal, allowing for finer control over the trade-off between unlearning efficacy and generative fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Framework for Diffusion Model Unlearning with f-Divergence
Novello, Nicola
Fontana, Federico
Cinque, Luigi
Gunduz, Deniz
Tonello, Andrea M.
Machine Learning
Computer Vision and Pattern Recognition
Most existing methods for concept unlearning in text-to-image diffusion models minimize a mean squared error (MSE) loss between the denoiser outputs conditioned on a target and an anchor concept, which is implicitly the KL divergence between two Gaussians. We generalize this objective to any $f$-divergence, recovering MSE as the KL instance, and identify a family of $α$-divergences whose Gaussian closed-form yields cheap, MSE-like training objectives. For the remaining $f$-divergences, we provide a min-max objective based on the variational formulation of the $f$-divergence. We theoretically analyze and numerically validate how different $f$-divergences impact the gradient magnitude and the convergence properties of the algorithm, affecting the quality of unlearning. For instance, we observe that the Hellinger closed-form instance consistently dominates MSE across multiple scenarios. More generally, the proposed unified framework offers a flexible paradigm for selecting the optimal divergence based on the application and user goal, allowing for finer control over the trade-off between unlearning efficacy and generative fidelity.
title A Unified Framework for Diffusion Model Unlearning with f-Divergence
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.21167