REACT: Revealing Evolutionary Action Consequence Trajectories for Interpretable Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Altmann, Philipp, Davignon, Céline, Zorn, Maximilian, Ritz, Fabian, Linnhoff-Popien, Claudia, Gabor, Thomas
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914740963377152
author Altmann, Philipp
Davignon, Céline
Zorn, Maximilian
Ritz, Fabian
Linnhoff-Popien, Claudia
Gabor, Thomas
author_facet Altmann, Philipp
Davignon, Céline
Zorn, Maximilian
Ritz, Fabian
Linnhoff-Popien, Claudia
Gabor, Thomas
contents To enhance the interpretability of Reinforcement Learning (RL), we propose Revealing Evolutionary Action Consequence Trajectories (REACT). In contrast to the prevalent practice of validating RL models based on their optimal behavior learned during training, we posit that considering a range of edge-case trajectories provides a more comprehensive understanding of their inherent behavior. To induce such scenarios, we introduce a disturbance to the initial state, optimizing it through an evolutionary algorithm to generate a diverse population of demonstrations. To evaluate the fitness of trajectories, REACT incorporates a joint fitness function that encourages both local and global diversity in the encountered states and chosen actions. Through assessments with policies trained for varying durations in discrete and continuous environments, we demonstrate the descriptive power of REACT. Our results highlight its effectiveness in revealing nuanced aspects of RL models' behavior beyond optimal performance, thereby contributing to improved interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03359
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle REACT: Revealing Evolutionary Action Consequence Trajectories for Interpretable Reinforcement Learning
Altmann, Philipp
Davignon, Céline
Zorn, Maximilian
Ritz, Fabian
Linnhoff-Popien, Claudia
Gabor, Thomas
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
To enhance the interpretability of Reinforcement Learning (RL), we propose Revealing Evolutionary Action Consequence Trajectories (REACT). In contrast to the prevalent practice of validating RL models based on their optimal behavior learned during training, we posit that considering a range of edge-case trajectories provides a more comprehensive understanding of their inherent behavior. To induce such scenarios, we introduce a disturbance to the initial state, optimizing it through an evolutionary algorithm to generate a diverse population of demonstrations. To evaluate the fitness of trajectories, REACT incorporates a joint fitness function that encourages both local and global diversity in the encountered states and chosen actions. Through assessments with policies trained for varying durations in discrete and continuous environments, we demonstrate the descriptive power of REACT. Our results highlight its effectiveness in revealing nuanced aspects of RL models' behavior beyond optimal performance, thereby contributing to improved interpretability.
title REACT: Revealing Evolutionary Action Consequence Trajectories for Interpretable Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2404.03359