FACE: Faithful Automatic Concept Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhusal, Dipkamal, Clifford, Michael, Rampazzi, Sara, Rastogi, Nidhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918159671361536
author Bhusal, Dipkamal
Clifford, Michael
Rampazzi, Sara
Rastogi, Nidhi
author_facet Bhusal, Dipkamal
Clifford, Michael
Rampazzi, Sara
Rastogi, Nidhi
contents Interpreting deep neural networks through concept-based explanations offers a bridge between low-level features and high-level human-understandable semantics. However, existing automatic concept discovery methods often fail to align these extracted concepts with the model's true decision-making process, thereby compromising explanation faithfulness. In this work, we propose FACE (Faithful Automatic Concept Extraction), a novel framework that augments Non-negative Matrix Factorization (NMF) with a Kullback-Leibler (KL) divergence regularization term to ensure alignment between the model's original and concept-based predictions. Unlike prior methods that operate solely on encoder activations, FACE incorporates classifier supervision during concept learning, enforcing predictive consistency and enabling faithful explanations. We provide theoretical guarantees showing that minimizing the KL divergence bounds the deviation in predictive distributions, thereby promoting faithful local linearity in the learned concept space. Systematic evaluations on ImageNet, COCO, and CelebA datasets demonstrate that FACE outperforms existing methods across faithfulness and sparsity metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FACE: Faithful Automatic Concept Extraction
Bhusal, Dipkamal
Clifford, Michael
Rampazzi, Sara
Rastogi, Nidhi
Computer Vision and Pattern Recognition
Artificial Intelligence
Interpreting deep neural networks through concept-based explanations offers a bridge between low-level features and high-level human-understandable semantics. However, existing automatic concept discovery methods often fail to align these extracted concepts with the model's true decision-making process, thereby compromising explanation faithfulness. In this work, we propose FACE (Faithful Automatic Concept Extraction), a novel framework that augments Non-negative Matrix Factorization (NMF) with a Kullback-Leibler (KL) divergence regularization term to ensure alignment between the model's original and concept-based predictions. Unlike prior methods that operate solely on encoder activations, FACE incorporates classifier supervision during concept learning, enforcing predictive consistency and enabling faithful explanations. We provide theoretical guarantees showing that minimizing the KL divergence bounds the deviation in predictive distributions, thereby promoting faithful local linearity in the learned concept space. Systematic evaluations on ImageNet, COCO, and CelebA datasets demonstrate that FACE outperforms existing methods across faithfulness and sparsity metrics.
title FACE: Faithful Automatic Concept Extraction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.11675