From Visual Explanations to Counterfactual Explanations with Latent Diffusion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luu, Tung, Le, Nam, Le, Duc, Le, Bac
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913790823497728
author Luu, Tung
Le, Nam
Le, Duc
Le, Bac
author_facet Luu, Tung
Le, Nam
Le, Duc
Le, Bac
contents Visual counterfactual explanations are ideal hypothetical images that change the decision-making of the classifier with high confidence toward the desired class while remaining visually plausible and close to the initial image. In this paper, we propose a new approach to tackle two key challenges in recent prominent works: i) determining which specific counterfactual features are crucial for distinguishing the "concept" of the target class from the original class, and ii) supplying valuable explanations for the non-robust classifier without relying on the support of an adversarially robust model. Our method identifies the essential region for modification through algorithms that provide visual explanations, and then our framework generates realistic counterfactual explanations by combining adversarial attacks based on pruning the adversarial gradient of the target classifier and the latent diffusion model. The proposed method outperforms previous state-of-the-art results on various evaluation criteria on ImageNet and CelebA-HQ datasets. In general, our method can be applied to arbitrary classifiers, highlight the strong association between visual and counterfactual explanations, make semantically meaningful changes from the target classifier, and provide observers with subtle counterfactual images.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Visual Explanations to Counterfactual Explanations with Latent Diffusion
Luu, Tung
Le, Nam
Le, Duc
Le, Bac
Computer Vision and Pattern Recognition
Visual counterfactual explanations are ideal hypothetical images that change the decision-making of the classifier with high confidence toward the desired class while remaining visually plausible and close to the initial image. In this paper, we propose a new approach to tackle two key challenges in recent prominent works: i) determining which specific counterfactual features are crucial for distinguishing the "concept" of the target class from the original class, and ii) supplying valuable explanations for the non-robust classifier without relying on the support of an adversarially robust model. Our method identifies the essential region for modification through algorithms that provide visual explanations, and then our framework generates realistic counterfactual explanations by combining adversarial attacks based on pruning the adversarial gradient of the target classifier and the latent diffusion model. The proposed method outperforms previous state-of-the-art results on various evaluation criteria on ImageNet and CelebA-HQ datasets. In general, our method can be applied to arbitrary classifiers, highlight the strong association between visual and counterfactual explanations, make semantically meaningful changes from the target classifier, and provide observers with subtle counterfactual images.
title From Visual Explanations to Counterfactual Explanations with Latent Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.09202