Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Koma, Arian Komaei, Kasaei, Seyed Amir, Aghayari, Ali, Sadeghzadeh, AmirMahdi, Rohban, Mohammad Hossein
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914448498753536
author Koma, Arian Komaei
Kasaei, Seyed Amir
Aghayari, Ali
Sadeghzadeh, AmirMahdi
Rohban, Mohammad Hossein
author_facet Koma, Arian Komaei
Kasaei, Seyed Amir
Aghayari, Ali
Sadeghzadeh, AmirMahdi
Rohban, Mohammad Hossein
contents Post-hoc unlearning has emerged as a practical mechanism for removing undesirable concepts from large text-to-image diffusion models. However, prior work primarily evaluates unlearning through erasure success; its impact on broader generative capabilities remains poorly understood. In this work, we conduct a systematic empirical study of concept unlearning through the lens of compositional text-to-image generation. Focusing on nudity removal in Stable Diffusion 1.4, we evaluate a diverse set of state-of-the-art unlearning methods using T2I-CompBench++ and GenEval, alongside established unlearning benchmarks. Our results reveal a consistent trade-off between unlearning effectiveness and compositional integrity: methods that achieve strong erasure frequently incur substantial degradation in attribute binding, spatial reasoning, and counting. Conversely, approaches that preserve compositional structure often fail to provide robust erasure. These findings highlight limitations of current evaluation practices and underscore the need for unlearning objectives that explicitly account for semantic preservation beyond targeted suppression.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04575
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models
Koma, Arian Komaei
Kasaei, Seyed Amir
Aghayari, Ali
Sadeghzadeh, AmirMahdi
Rohban, Mohammad Hossein
Computer Vision and Pattern Recognition
Post-hoc unlearning has emerged as a practical mechanism for removing undesirable concepts from large text-to-image diffusion models. However, prior work primarily evaluates unlearning through erasure success; its impact on broader generative capabilities remains poorly understood. In this work, we conduct a systematic empirical study of concept unlearning through the lens of compositional text-to-image generation. Focusing on nudity removal in Stable Diffusion 1.4, we evaluate a diverse set of state-of-the-art unlearning methods using T2I-CompBench++ and GenEval, alongside established unlearning benchmarks. Our results reveal a consistent trade-off between unlearning effectiveness and compositional integrity: methods that achieve strong erasure frequently incur substantial degradation in attribute binding, spatial reasoning, and counting. Conversely, approaches that preserve compositional structure often fail to provide robust erasure. These findings highlight limitations of current evaluation practices and underscore the need for unlearning objectives that explicitly account for semantic preservation beyond targeted suppression.
title Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.04575