Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pahde, Frederik, Dreyer, Maximilian, Weber, Leander, Weckbecker, Moritz, Anders, Christopher J., Wiegand, Thomas, Samek, Wojciech, Lapuschkin, Sebastian
Format: Preprint
Publié: 2022
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916723808010240
author Pahde, Frederik
Dreyer, Maximilian
Weber, Leander
Weckbecker, Moritz
Anders, Christopher J.
Wiegand, Thomas
Samek, Wojciech
Lapuschkin, Sebastian
author_facet Pahde, Frederik
Dreyer, Maximilian
Weber, Leander
Weckbecker, Moritz
Anders, Christopher J.
Wiegand, Thomas
Samek, Wojciech
Lapuschkin, Sebastian
contents With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CAVs are computed by leveraging linear classifiers optimizing the separability of latent representations of samples with and without a given concept. However, in this paper we show that such a separability-oriented computation leads to solutions, which may diverge from the actual goal of precisely modeling the concept direction. This discrepancy can be attributed to the significant influence of distractor directions, i.e., signals unrelated to the concept, which are picked up by filters (i.e., weights) of linear models to optimize class-separability. To address this, we introduce pattern-based CAVs, solely focussing on concept signals, thereby providing more accurate concept directions. We evaluate various CAV methods in terms of their alignment with the true concept direction and their impact on CAV applications, including concept sensitivity testing and model correction for shortcut behavior caused by data artifacts. We demonstrate the benefits of pattern-based CAVs using the Pediatric Bone Age, ISIC2019, and FunnyBirds datasets with VGG, ResNet, ReXNet, EfficientNet, and Vision Transformer as model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2202_03482
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
Pahde, Frederik
Dreyer, Maximilian
Weber, Leander
Weckbecker, Moritz
Anders, Christopher J.
Wiegand, Thomas
Samek, Wojciech
Lapuschkin, Sebastian
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CAVs are computed by leveraging linear classifiers optimizing the separability of latent representations of samples with and without a given concept. However, in this paper we show that such a separability-oriented computation leads to solutions, which may diverge from the actual goal of precisely modeling the concept direction. This discrepancy can be attributed to the significant influence of distractor directions, i.e., signals unrelated to the concept, which are picked up by filters (i.e., weights) of linear models to optimize class-separability. To address this, we introduce pattern-based CAVs, solely focussing on concept signals, thereby providing more accurate concept directions. We evaluate various CAV methods in terms of their alignment with the true concept direction and their impact on CAV applications, including concept sensitivity testing and model correction for shortcut behavior caused by data artifacts. We demonstrate the benefits of pattern-based CAVs using the Pediatric Bone Age, ISIC2019, and FunnyBirds datasets with VGG, ResNet, ReXNet, EfficientNet, and Vision Transformer as model architectures.
title Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2202.03482