Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Septon, Yael, Huber, Tobias, André, Elisabeth, Amir, Ofra
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917598394843136
author Septon, Yael
Huber, Tobias
André, Elisabeth
Amir, Ofra
author_facet Septon, Yael
Huber, Tobias
André, Elisabeth
Amir, Ofra
contents Explaining the behavior of reinforcement learning agents operating in sequential decision-making settings is challenging, as their behavior is affected by a dynamic environment and delayed rewards. Methods that help users understand the behavior of such agents can roughly be divided into local explanations that analyze specific decisions of the agents and global explanations that convey the general strategy of the agents. In this work, we study a novel combination of local and global explanations for reinforcement learning agents. Specifically, we combine reward decomposition, a local explanation method that exposes which components of the reward function influenced a specific decision, and HIGHLIGHTS, a global explanation method that shows a summary of the agent's behavior in decisive states. We conducted two user studies to evaluate the integration of these explanation methods and their respective benefits. Our results show significant benefits for both methods. In general, we found that the local reward decomposition was more useful for identifying the agents' priorities. However, when there was only a minor difference between the agents' preferences, then the global information provided by HIGHLIGHTS additionally improved participants' understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2210_11825
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents
Septon, Yael
Huber, Tobias
André, Elisabeth
Amir, Ofra
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Explaining the behavior of reinforcement learning agents operating in sequential decision-making settings is challenging, as their behavior is affected by a dynamic environment and delayed rewards. Methods that help users understand the behavior of such agents can roughly be divided into local explanations that analyze specific decisions of the agents and global explanations that convey the general strategy of the agents. In this work, we study a novel combination of local and global explanations for reinforcement learning agents. Specifically, we combine reward decomposition, a local explanation method that exposes which components of the reward function influenced a specific decision, and HIGHLIGHTS, a global explanation method that shows a summary of the agent's behavior in decisive states. We conducted two user studies to evaluate the integration of these explanation methods and their respective benefits. Our results show significant benefits for both methods. In general, we found that the local reward decomposition was more useful for identifying the agents' priorities. However, when there was only a minor difference between the agents' preferences, then the global information provided by HIGHLIGHTS additionally improved participants' understanding.
title Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2210.11825