Intrinsic User-Centric Interpretability through Global Mixture of Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Swamy, Vinitra, Montariol, Syrielle, Blackwell, Julian, Frej, Jibril, Jaggi, Martin, Käser, Tanja
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909625604898816
author Swamy, Vinitra
Montariol, Syrielle
Blackwell, Julian
Frej, Jibril
Jaggi, Martin
Käser, Tanja
author_facet Swamy, Vinitra
Montariol, Syrielle
Blackwell, Julian
Frej, Jibril
Jaggi, Martin
Käser, Tanja
contents In human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption. Towards these two goals, intrinsically interpretable deep learning models have gained popularity, focusing on accurate predictions alongside faithful explanations. However, there exists a gap in the human-centeredness of these approaches, which often produce nuanced and complex explanations that are not easily actionable for downstream users. We present InterpretCC (interpretable conditional computation), a family of intrinsically interpretable neural networks at a unique point in the design space that optimizes for ease of human understanding and explanation faithfulness, while maintaining comparable performance to state-of-the-art models. InterpretCC achieves this through adaptive sparse activation of features before prediction, allowing the model to use a different, minimal set of features for each instance. We extend this idea into an interpretable, global mixture-of-experts (MoE) model that allows users to specify topics of interest, discretely separates the feature space for each data point into topical subnetworks, and adaptively and sparsely activates these topical subnetworks for prediction. We apply InterpretCC for text, time series and tabular data across several real-world datasets, demonstrating comparable performance with non-interpretable baselines and outperforming intrinsically interpretable baselines. Through a user study involving 56 teachers, InterpretCC explanations are found to have higher actionability and usefulness over other intrinsically interpretable approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02933
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Intrinsic User-Centric Interpretability through Global Mixture of Experts
Swamy, Vinitra
Montariol, Syrielle
Blackwell, Julian
Frej, Jibril
Jaggi, Martin
Käser, Tanja
Machine Learning
Computers and Society
Human-Computer Interaction
In human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption. Towards these two goals, intrinsically interpretable deep learning models have gained popularity, focusing on accurate predictions alongside faithful explanations. However, there exists a gap in the human-centeredness of these approaches, which often produce nuanced and complex explanations that are not easily actionable for downstream users. We present InterpretCC (interpretable conditional computation), a family of intrinsically interpretable neural networks at a unique point in the design space that optimizes for ease of human understanding and explanation faithfulness, while maintaining comparable performance to state-of-the-art models. InterpretCC achieves this through adaptive sparse activation of features before prediction, allowing the model to use a different, minimal set of features for each instance. We extend this idea into an interpretable, global mixture-of-experts (MoE) model that allows users to specify topics of interest, discretely separates the feature space for each data point into topical subnetworks, and adaptively and sparsely activates these topical subnetworks for prediction. We apply InterpretCC for text, time series and tabular data across several real-world datasets, demonstrating comparable performance with non-interpretable baselines and outperforming intrinsically interpretable baselines. Through a user study involving 56 teachers, InterpretCC explanations are found to have higher actionability and usefulness over other intrinsically interpretable approaches.
title Intrinsic User-Centric Interpretability through Global Mixture of Experts
topic Machine Learning
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2402.02933