Linking Model Intervention to Causal Interpretation in Model Explanation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Debo, Xu, Ziqi, Li, Jiuyong, Liu, Lin, Yu, Kui, Le, Thuc Duy, Liu, Jixue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917810305761280
author Cheng, Debo
Xu, Ziqi
Li, Jiuyong
Liu, Lin
Yu, Kui
Le, Thuc Duy
Liu, Jixue
author_facet Cheng, Debo
Xu, Ziqi
Li, Jiuyong
Liu, Lin
Yu, Kui
Le, Thuc Duy
Liu, Jixue
contents Intervention intuition is often used in model explanation where the intervention effect of a feature on the outcome is quantified by the difference of a model prediction when the feature value is changed from the current value to the baseline value. Such a model intervention effect of a feature is inherently association. In this paper, we will study the conditions when an intuitive model intervention effect has a causal interpretation, i.e., when it indicates whether a feature is a direct cause of the outcome. This work links the model intervention effect to the causal interpretation of a model. Such an interpretation capability is important since it indicates whether a machine learning model is trustworthy to domain experts. The conditions also reveal the limitations of using a model intervention effect for causal interpretation in an environment with unobserved features. Experiments on semi-synthetic datasets have been conducted to validate theorems and show the potential for using the model intervention effect for model interpretation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15648
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Linking Model Intervention to Causal Interpretation in Model Explanation
Cheng, Debo
Xu, Ziqi
Li, Jiuyong
Liu, Lin
Yu, Kui
Le, Thuc Duy
Liu, Jixue
Machine Learning
Methodology
Intervention intuition is often used in model explanation where the intervention effect of a feature on the outcome is quantified by the difference of a model prediction when the feature value is changed from the current value to the baseline value. Such a model intervention effect of a feature is inherently association. In this paper, we will study the conditions when an intuitive model intervention effect has a causal interpretation, i.e., when it indicates whether a feature is a direct cause of the outcome. This work links the model intervention effect to the causal interpretation of a model. Such an interpretation capability is important since it indicates whether a machine learning model is trustworthy to domain experts. The conditions also reveal the limitations of using a model intervention effect for causal interpretation in an environment with unobserved features. Experiments on semi-synthetic datasets have been conducted to validate theorems and show the potential for using the model intervention effect for model interpretation.
title Linking Model Intervention to Causal Interpretation in Model Explanation
topic Machine Learning
Methodology
url https://arxiv.org/abs/2410.15648