Linking Model Intervention to Causal Interpretation in Model Explanation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917810305761280 |
|---|---|
| author | Cheng, Debo Xu, Ziqi Li, Jiuyong Liu, Lin Yu, Kui Le, Thuc Duy Liu, Jixue |
| author_facet | Cheng, Debo Xu, Ziqi Li, Jiuyong Liu, Lin Yu, Kui Le, Thuc Duy Liu, Jixue |
| contents | Intervention intuition is often used in model explanation where the intervention effect of a feature on the outcome is quantified by the difference of a model prediction when the feature value is changed from the current value to the baseline value. Such a model intervention effect of a feature is inherently association. In this paper, we will study the conditions when an intuitive model intervention effect has a causal interpretation, i.e., when it indicates whether a feature is a direct cause of the outcome. This work links the model intervention effect to the causal interpretation of a model. Such an interpretation capability is important since it indicates whether a machine learning model is trustworthy to domain experts. The conditions also reveal the limitations of using a model intervention effect for causal interpretation in an environment with unobserved features. Experiments on semi-synthetic datasets have been conducted to validate theorems and show the potential for using the model intervention effect for model interpretation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_15648 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Linking Model Intervention to Causal Interpretation in Model Explanation Cheng, Debo Xu, Ziqi Li, Jiuyong Liu, Lin Yu, Kui Le, Thuc Duy Liu, Jixue Machine Learning Methodology Intervention intuition is often used in model explanation where the intervention effect of a feature on the outcome is quantified by the difference of a model prediction when the feature value is changed from the current value to the baseline value. Such a model intervention effect of a feature is inherently association. In this paper, we will study the conditions when an intuitive model intervention effect has a causal interpretation, i.e., when it indicates whether a feature is a direct cause of the outcome. This work links the model intervention effect to the causal interpretation of a model. Such an interpretation capability is important since it indicates whether a machine learning model is trustworthy to domain experts. The conditions also reveal the limitations of using a model intervention effect for causal interpretation in an environment with unobserved features. Experiments on semi-synthetic datasets have been conducted to validate theorems and show the potential for using the model intervention effect for model interpretation. |
| title | Linking Model Intervention to Causal Interpretation in Model Explanation |
| topic | Machine Learning Methodology |
| url | https://arxiv.org/abs/2410.15648 |