Rule-Guided Reinforcement Learning Policy Evaluation and Improvement

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tappler, Martin, Lopez-Miguel, Ignacio D., Tschiatschek, Sebastian, Bartocci, Ezio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910871844814848
author Tappler, Martin
Lopez-Miguel, Ignacio D.
Tschiatschek, Sebastian
Bartocci, Ezio
author_facet Tappler, Martin
Lopez-Miguel, Ignacio D.
Tschiatschek, Sebastian
Bartocci, Ezio
contents We consider the challenging problem of using domain knowledge to improve deep reinforcement learning policies. To this end, we propose LEGIBLE, a novel approach, following a multi-step process, which starts by mining rules from a deep RL policy, constituting a partially symbolic representation. These rules describe which decisions the RL policy makes and which it avoids making. In the second step, we generalize the mined rules using domain knowledge expressed as metamorphic relations. We adapt these relations from software testing to RL to specify expected changes of actions in response to changes in observations. The third step is evaluating generalized rules to determine which generalizations improve performance when enforced. These improvements show weaknesses in the policy, where it has not learned the general rules and thus can be improved by rule guidance. LEGIBLE supported by metamorphic relations provides a principled way of expressing and enforcing domain knowledge about RL environments. We show the efficacy of our approach by demonstrating that it effectively finds weaknesses, accompanied by explanations of these weaknesses, in eleven RL environments and by showcasing that guiding policy execution with rules improves performance w.r.t. gained reward.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rule-Guided Reinforcement Learning Policy Evaluation and Improvement
Tappler, Martin
Lopez-Miguel, Ignacio D.
Tschiatschek, Sebastian
Bartocci, Ezio
Machine Learning
Software Engineering
We consider the challenging problem of using domain knowledge to improve deep reinforcement learning policies. To this end, we propose LEGIBLE, a novel approach, following a multi-step process, which starts by mining rules from a deep RL policy, constituting a partially symbolic representation. These rules describe which decisions the RL policy makes and which it avoids making. In the second step, we generalize the mined rules using domain knowledge expressed as metamorphic relations. We adapt these relations from software testing to RL to specify expected changes of actions in response to changes in observations. The third step is evaluating generalized rules to determine which generalizations improve performance when enforced. These improvements show weaknesses in the policy, where it has not learned the general rules and thus can be improved by rule guidance. LEGIBLE supported by metamorphic relations provides a principled way of expressing and enforcing domain knowledge about RL environments. We show the efficacy of our approach by demonstrating that it effectively finds weaknesses, accompanied by explanations of these weaknesses, in eleven RL environments and by showcasing that guiding policy execution with rules improves performance w.r.t. gained reward.
title Rule-Guided Reinforcement Learning Policy Evaluation and Improvement
topic Machine Learning
Software Engineering
url https://arxiv.org/abs/2503.09270