A Framework for Causal Concept-based Model Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bjøru, Anna Rodum, Lysnæs-Larsen, Jacob, Jørgensen, Oskar, Strümke, Inga, Langseth, Helge
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918228352040960
author Bjøru, Anna Rodum
Lysnæs-Larsen, Jacob
Jørgensen, Oskar
Strümke, Inga
Langseth, Helge
author_facet Bjøru, Anna Rodum
Lysnæs-Larsen, Jacob
Jørgensen, Oskar
Strümke, Inga
Langseth, Helge
contents This work presents a conceptual framework for causal concept-based post-hoc Explainable Artificial Intelligence (XAI), based on the requirements that explanations for non-interpretable models should be understandable as well as faithful to the model being explained. Local and global explanations are generated by calculating the probability of sufficiency of concept interventions. Example explanations are presented, generated with a proof-of-concept model made to explain classifiers trained on the CelebA dataset. Understandability is demonstrated through a clear concept-based vocabulary, subject to an implicit causal interpretation. Fidelity is addressed by highlighting important framework assumptions, stressing that the context of explanation interpretation must align with the context of explanation generation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Framework for Causal Concept-based Model Explanations
Bjøru, Anna Rodum
Lysnæs-Larsen, Jacob
Jørgensen, Oskar
Strümke, Inga
Langseth, Helge
Artificial Intelligence
This work presents a conceptual framework for causal concept-based post-hoc Explainable Artificial Intelligence (XAI), based on the requirements that explanations for non-interpretable models should be understandable as well as faithful to the model being explained. Local and global explanations are generated by calculating the probability of sufficiency of concept interventions. Example explanations are presented, generated with a proof-of-concept model made to explain classifiers trained on the CelebA dataset. Understandability is demonstrated through a clear concept-based vocabulary, subject to an implicit causal interpretation. Fidelity is addressed by highlighting important framework assumptions, stressing that the context of explanation interpretation must align with the context of explanation generation.
title A Framework for Causal Concept-based Model Explanations
topic Artificial Intelligence
url https://arxiv.org/abs/2512.02735