Explainable Reinforcement Learning Agents Using World Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Singh, Madhuri, Alabdulkarim, Amal, Mansi, Gennie, Riedl, Mark O.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918126251147264
author Singh, Madhuri
Alabdulkarim, Amal
Mansi, Gennie
Riedl, Mark O.
author_facet Singh, Madhuri
Alabdulkarim, Amal
Mansi, Gennie
Riedl, Mark O.
contents Explainable AI (XAI) systems have been proposed to help people understand how AI systems produce outputs and behaviors. Explainable Reinforcement Learning (XRL) has an added complexity due to the temporal nature of sequential decision-making. Further, non-AI experts do not necessarily have the ability to alter an agent or its policy. We introduce a technique for using World Models to generate explanations for Model-Based Deep RL agents. World Models predict how the world will change when actions are performed, allowing for the generation of counterfactual trajectories. However, identifying what a user wanted the agent to do is not enough to understand why the agent did something else. We augment Model-Based RL agents with a Reverse World Model, which predicts what the state of the world should have been for the agent to prefer a given counterfactual action. We show that explanations that show users what the world should have been like significantly increase their understanding of the agent policy. We hypothesize that our explanations can help users learn how to control the agents execution through by manipulating the environment.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08073
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explainable Reinforcement Learning Agents Using World Models
Singh, Madhuri
Alabdulkarim, Amal
Mansi, Gennie
Riedl, Mark O.
Artificial Intelligence
Explainable AI (XAI) systems have been proposed to help people understand how AI systems produce outputs and behaviors. Explainable Reinforcement Learning (XRL) has an added complexity due to the temporal nature of sequential decision-making. Further, non-AI experts do not necessarily have the ability to alter an agent or its policy. We introduce a technique for using World Models to generate explanations for Model-Based Deep RL agents. World Models predict how the world will change when actions are performed, allowing for the generation of counterfactual trajectories. However, identifying what a user wanted the agent to do is not enough to understand why the agent did something else. We augment Model-Based RL agents with a Reverse World Model, which predicts what the state of the world should have been for the agent to prefer a given counterfactual action. We show that explanations that show users what the world should have been like significantly increase their understanding of the agent policy. We hypothesize that our explanations can help users learn how to control the agents execution through by manipulating the environment.
title Explainable Reinforcement Learning Agents Using World Models
topic Artificial Intelligence
url https://arxiv.org/abs/2505.08073