S-CFE: Simple Counterfactual Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sadiku, Shpresim, Wagner, Moritz, Nagarajan, Sai Ganesh, Pokutta, Sebastian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917077006155776
author Sadiku, Shpresim
Wagner, Moritz
Nagarajan, Sai Ganesh
Pokutta, Sebastian
author_facet Sadiku, Shpresim
Wagner, Moritz
Nagarajan, Sai Ganesh
Pokutta, Sebastian
contents We study the problem of finding optimal sparse, manifold-aligned counterfactual explanations for classifiers. Canonically, this can be formulated as an optimization problem with multiple non-convex components, including classifier loss functions and manifold alignment (or \emph{plausibility}) metrics. The added complexity of enforcing \emph{sparsity}, or shorter explanations, complicates the problem further. Existing methods often focus on specific models and plausibility measures, relying on convex $\ell_1$ regularizers to enforce sparsity. In this paper, we tackle the canonical formulation using the accelerated proximal gradient (APG) method, a simple yet efficient first-order procedure capable of handling smooth non-convex objectives and non-smooth $\ell_p$ (where $0 \leq p < 1$) regularizers. This enables our approach to seamlessly incorporate various classifiers and plausibility measures while producing sparser solutions. Our algorithm only requires differentiable data-manifold regularizers and supports box constraints for bounded feature ranges, ensuring the generated counterfactuals remain \emph{actionable}. Finally, experiments on real-world datasets demonstrate that our approach effectively produces sparse, manifold-aligned counterfactual explanations while maintaining proximity to the factual data and computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15723
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle S-CFE: Simple Counterfactual Explanations
Sadiku, Shpresim
Wagner, Moritz
Nagarajan, Sai Ganesh
Pokutta, Sebastian
Machine Learning
Optimization and Control
We study the problem of finding optimal sparse, manifold-aligned counterfactual explanations for classifiers. Canonically, this can be formulated as an optimization problem with multiple non-convex components, including classifier loss functions and manifold alignment (or \emph{plausibility}) metrics. The added complexity of enforcing \emph{sparsity}, or shorter explanations, complicates the problem further. Existing methods often focus on specific models and plausibility measures, relying on convex $\ell_1$ regularizers to enforce sparsity. In this paper, we tackle the canonical formulation using the accelerated proximal gradient (APG) method, a simple yet efficient first-order procedure capable of handling smooth non-convex objectives and non-smooth $\ell_p$ (where $0 \leq p < 1$) regularizers. This enables our approach to seamlessly incorporate various classifiers and plausibility measures while producing sparser solutions. Our algorithm only requires differentiable data-manifold regularizers and supports box constraints for bounded feature ranges, ensuring the generated counterfactuals remain \emph{actionable}. Finally, experiments on real-world datasets demonstrate that our approach effectively produces sparse, manifold-aligned counterfactual explanations while maintaining proximity to the factual data and computational efficiency.
title S-CFE: Simple Counterfactual Explanations
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2410.15723