Understanding Inter-Concept Relationships in Concept-Based Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raman, Naveen, Zarlenga, Mateo Espinosa, Jamnik, Mateja
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929362354307072
author Raman, Naveen
Zarlenga, Mateo Espinosa
Jamnik, Mateja
author_facet Raman, Naveen
Zarlenga, Mateo Espinosa
Jamnik, Mateja
contents Concept-based explainability methods provide insight into deep learning systems by constructing explanations using human-understandable concepts. While the literature on human reasoning demonstrates that we exploit relationships between concepts when solving tasks, it is unclear whether concept-based methods incorporate the rich structure of inter-concept relationships. We analyse the concept representations learnt by concept-based models to understand whether these models correctly capture inter-concept relationships. First, we empirically demonstrate that state-of-the-art concept-based models produce representations that lack stability and robustness, and such methods fail to capture inter-concept relationships. Then, we develop a novel algorithm which leverages inter-concept relationships to improve concept intervention accuracy, demonstrating how correctly capturing inter-concept relationships can improve downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18217
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Inter-Concept Relationships in Concept-Based Models
Raman, Naveen
Zarlenga, Mateo Espinosa
Jamnik, Mateja
Machine Learning
Concept-based explainability methods provide insight into deep learning systems by constructing explanations using human-understandable concepts. While the literature on human reasoning demonstrates that we exploit relationships between concepts when solving tasks, it is unclear whether concept-based methods incorporate the rich structure of inter-concept relationships. We analyse the concept representations learnt by concept-based models to understand whether these models correctly capture inter-concept relationships. First, we empirically demonstrate that state-of-the-art concept-based models produce representations that lack stability and robustness, and such methods fail to capture inter-concept relationships. Then, we develop a novel algorithm which leverages inter-concept relationships to improve concept intervention accuracy, demonstrating how correctly capturing inter-concept relationships can improve downstream tasks.
title Understanding Inter-Concept Relationships in Concept-Based Models
topic Machine Learning
url https://arxiv.org/abs/2405.18217