Explaining Reinforcement Learning: A Counterfactual Shapley Values Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Yiwei, Zhang, Qi, McAreavey, Kevin, Liu, Weiru
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914902540550144
author Shi, Yiwei
Zhang, Qi
McAreavey, Kevin
Liu, Weiru
author_facet Shi, Yiwei
Zhang, Qi
McAreavey, Kevin
Liu, Weiru
contents This paper introduces a novel approach Counterfactual Shapley Values (CSV), which enhances explainability in reinforcement learning (RL) by integrating counterfactual analysis with Shapley Values. The approach aims to quantify and compare the contributions of different state dimensions to various action choices. To more accurately analyze these impacts, we introduce new characteristic value functions, the ``Counterfactual Difference Characteristic Value" and the ``Average Counterfactual Difference Characteristic Value." These functions help calculate the Shapley values to evaluate the differences in contributions between optimal and non-optimal actions. Experiments across several RL domains, such as GridWorld, FrozenLake, and Taxi, demonstrate the effectiveness of the CSV method. The results show that this method not only improves transparency in complex RL systems but also quantifies the differences across various decisions.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02529
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Explaining Reinforcement Learning: A Counterfactual Shapley Values Approach
Shi, Yiwei
Zhang, Qi
McAreavey, Kevin
Liu, Weiru
Artificial Intelligence
This paper introduces a novel approach Counterfactual Shapley Values (CSV), which enhances explainability in reinforcement learning (RL) by integrating counterfactual analysis with Shapley Values. The approach aims to quantify and compare the contributions of different state dimensions to various action choices. To more accurately analyze these impacts, we introduce new characteristic value functions, the ``Counterfactual Difference Characteristic Value" and the ``Average Counterfactual Difference Characteristic Value." These functions help calculate the Shapley values to evaluate the differences in contributions between optimal and non-optimal actions. Experiments across several RL domains, such as GridWorld, FrozenLake, and Taxi, demonstrate the effectiveness of the CSV method. The results show that this method not only improves transparency in complex RL systems but also quantifies the differences across various decisions.
title Explaining Reinforcement Learning: A Counterfactual Shapley Values Approach
topic Artificial Intelligence
url https://arxiv.org/abs/2408.02529