Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kohler, Hector, Delfosse, Quentin, Akrour, Riad, Kersting, Kristian, Preux, Philippe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911885623820288
author Kohler, Hector
Delfosse, Quentin
Akrour, Riad
Kersting, Kristian
Preux, Philippe
author_facet Kohler, Hector
Delfosse, Quentin
Akrour, Riad
Kersting, Kristian
Preux, Philippe
contents Deep reinforcement learning agents are prone to goal misalignments. The black-box nature of their policies hinders the detection and correction of such misalignments, and the trust necessary for real-world deployment. So far, solutions learning interpretable policies are inefficient or require many human priors. We propose INTERPRETER, a fast distillation method producing INTerpretable Editable tRee Programs for ReinforcEmenT lEaRning. We empirically demonstrate that INTERPRETER compact tree programs match oracles across a diverse set of sequential decision tasks and evaluate the impact of our design choices on interpretability and performances. We show that our policies can be interpreted and edited to correct misalignments on Atari games and to explain real farming strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
Kohler, Hector
Delfosse, Quentin
Akrour, Riad
Kersting, Kristian
Preux, Philippe
Artificial Intelligence
Machine Learning
Deep reinforcement learning agents are prone to goal misalignments. The black-box nature of their policies hinders the detection and correction of such misalignments, and the trust necessary for real-world deployment. So far, solutions learning interpretable policies are inefficient or require many human priors. We propose INTERPRETER, a fast distillation method producing INTerpretable Editable tRee Programs for ReinforcEmenT lEaRning. We empirically demonstrate that INTERPRETER compact tree programs match oracles across a diverse set of sequential decision tasks and evaluate the impact of our design choices on interpretability and performances. We show that our policies can be interpreted and edited to correct misalignments on Atari games and to explain real farming strategies.
title Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.14956