OpenHEXAI: An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Jiaqi, Lai, Vivian, Zhang, Yiming, Chen, Chacha, Hamilton, Paul, Ljubenkov, Davor, Lakkaraju, Himabindu, Tan, Chenhao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909132694487040
author Ma, Jiaqi
Lai, Vivian
Zhang, Yiming
Chen, Chacha
Hamilton, Paul
Ljubenkov, Davor
Lakkaraju, Himabindu
Tan, Chenhao
author_facet Ma, Jiaqi
Lai, Vivian
Zhang, Yiming
Chen, Chacha
Hamilton, Paul
Ljubenkov, Davor
Lakkaraju, Himabindu
Tan, Chenhao
contents Recently, there has been a surge of explainable AI (XAI) methods driven by the need for understanding machine learning model behaviors in high-stakes scenarios. However, properly evaluating the effectiveness of the XAI methods inevitably requires the involvement of human subjects, and conducting human-centered benchmarks is challenging in a number of ways: designing and implementing user studies is complex; numerous design choices in the design space of user study lead to problems of reproducibility; and running user studies can be challenging and even daunting for machine learning researchers. To address these challenges, this paper presents OpenHEXAI, an open-source framework for human-centered evaluation of XAI methods. OpenHEXAI features (1) a collection of diverse benchmark datasets, pre-trained models, and post hoc explanation methods; (2) an easy-to-use web application for user study; (3) comprehensive evaluation metrics for the effectiveness of post hoc explanation methods in the context of human-AI decision making tasks; (4) best practice recommendations of experiment documentation; and (5) convenient tools for power analysis and cost estimation. OpenHEAXI is the first large-scale infrastructural effort to facilitate human-centered benchmarks of XAI methods. It simplifies the design and implementation of user studies for XAI methods, thus allowing researchers and practitioners to focus on the scientific questions. Additionally, it enhances reproducibility through standardized designs. Based on OpenHEXAI, we further conduct a systematic benchmark of four state-of-the-art post hoc explanation methods and compare their impacts on human-AI decision making tasks in terms of accuracy, fairness, as well as users' trust and understanding of the machine learning model.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05565
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OpenHEXAI: An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning
Ma, Jiaqi
Lai, Vivian
Zhang, Yiming
Chen, Chacha
Hamilton, Paul
Ljubenkov, Davor
Lakkaraju, Himabindu
Tan, Chenhao
Human-Computer Interaction
Artificial Intelligence
Recently, there has been a surge of explainable AI (XAI) methods driven by the need for understanding machine learning model behaviors in high-stakes scenarios. However, properly evaluating the effectiveness of the XAI methods inevitably requires the involvement of human subjects, and conducting human-centered benchmarks is challenging in a number of ways: designing and implementing user studies is complex; numerous design choices in the design space of user study lead to problems of reproducibility; and running user studies can be challenging and even daunting for machine learning researchers. To address these challenges, this paper presents OpenHEXAI, an open-source framework for human-centered evaluation of XAI methods. OpenHEXAI features (1) a collection of diverse benchmark datasets, pre-trained models, and post hoc explanation methods; (2) an easy-to-use web application for user study; (3) comprehensive evaluation metrics for the effectiveness of post hoc explanation methods in the context of human-AI decision making tasks; (4) best practice recommendations of experiment documentation; and (5) convenient tools for power analysis and cost estimation. OpenHEAXI is the first large-scale infrastructural effort to facilitate human-centered benchmarks of XAI methods. It simplifies the design and implementation of user studies for XAI methods, thus allowing researchers and practitioners to focus on the scientific questions. Additionally, it enhances reproducibility through standardized designs. Based on OpenHEXAI, we further conduct a systematic benchmark of four state-of-the-art post hoc explanation methods and compare their impacts on human-AI decision making tasks in terms of accuracy, fairness, as well as users' trust and understanding of the machine learning model.
title OpenHEXAI: An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2403.05565