Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gajcin, Jasmina, Dusparic, Ivana
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916118898147328
author Gajcin, Jasmina
Dusparic, Ivana
author_facet Gajcin, Jasmina
Dusparic, Ivana
contents While AI algorithms have shown remarkable success in various fields, their lack of transparency hinders their application to real-life tasks. Although explanations targeted at non-experts are necessary for user trust and human-AI collaboration, the majority of explanation methods for AI are focused on developers and expert users. Counterfactual explanations are local explanations that offer users advice on what can be changed in the input for the output of the black-box model to change. Counterfactuals are user-friendly and provide actionable advice for achieving the desired output from the AI system. While extensively researched in supervised learning, there are few methods applying them to reinforcement learning (RL). In this work, we explore the reasons for the underrepresentation of a powerful explanation method in RL. We start by reviewing the current work in counterfactual explanations in supervised learning. Additionally, we explore the differences between counterfactual explanations in supervised learning and RL and identify the main challenges that prevent the adoption of methods from supervised in reinforcement learning. Finally, we redefine counterfactuals for RL and propose research directions for implementing counterfactuals in RL.
format Preprint
id arxiv_https___arxiv_org_abs_2210_11846
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities
Gajcin, Jasmina
Dusparic, Ivana
Artificial Intelligence
While AI algorithms have shown remarkable success in various fields, their lack of transparency hinders their application to real-life tasks. Although explanations targeted at non-experts are necessary for user trust and human-AI collaboration, the majority of explanation methods for AI are focused on developers and expert users. Counterfactual explanations are local explanations that offer users advice on what can be changed in the input for the output of the black-box model to change. Counterfactuals are user-friendly and provide actionable advice for achieving the desired output from the AI system. While extensively researched in supervised learning, there are few methods applying them to reinforcement learning (RL). In this work, we explore the reasons for the underrepresentation of a powerful explanation method in RL. We start by reviewing the current work in counterfactual explanations in supervised learning. Additionally, we explore the differences between counterfactual explanations in supervised learning and RL and identify the main challenges that prevent the adoption of methods from supervised in reinforcement learning. Finally, we redefine counterfactuals for RL and propose research directions for implementing counterfactuals in RL.
title Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities
topic Artificial Intelligence
url https://arxiv.org/abs/2210.11846