Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
Fuente:
arXiv
Saved in:
| Main Authors: | Kohler, Hector, Delfosse, Quentin, Radji, Waris, Akrour, Riad, Preux, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024)
by: Kohler, Hector, et al.
Published: (2024)
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop
by: Kohler, Hector, et al.
Published: (2024)
by: Kohler, Hector, et al.
Published: (2024)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning
by: Berthelot, Yann, et al.
Published: (2026)
by: Berthelot, Yann, et al.
Published: (2026)
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
by: Driss, Brahim, et al.
Published: (2025)
by: Driss, Brahim, et al.
Published: (2025)
Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces
by: Kobanda, Anthony, et al.
Published: (2026)
by: Kobanda, Anthony, et al.
Published: (2026)
StaQ it! Growing neural networks for Policy Mirror Descent
by: Shilova, Alena, et al.
Published: (2025)
by: Shilova, Alena, et al.
Published: (2025)
Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning
by: Delfosse, Quentin, et al.
Published: (2024)
by: Delfosse, Quentin, et al.
Published: (2024)
GRAIL: Autonomous Concept Grounding for Neuro-Symbolic Reinforcement Learning
by: Shindo, Hikaru, et al.
Published: (2026)
by: Shindo, Hikaru, et al.
Published: (2026)
IDEQ: an improved diffusion model for the TSP
by: Basson, Mickael, et al.
Published: (2024)
by: Basson, Mickael, et al.
Published: (2024)
BlendRL: A Framework for Merging Symbolic and Neural Policy Learning
by: Shindo, Hikaru, et al.
Published: (2024)
by: Shindo, Hikaru, et al.
Published: (2024)
Deep Reinforcement Learning via Object-Centric Attention
by: Blüml, Jannis, et al.
Published: (2025)
by: Blüml, Jannis, et al.
Published: (2025)
Boosting deep Reinforcement Learning using pretraining with Logical Options
by: Ye, Zihan, et al.
Published: (2026)
by: Ye, Zihan, et al.
Published: (2026)
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
by: Delfosse, Quentin, et al.
Published: (2023)
by: Delfosse, Quentin, et al.
Published: (2023)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
by: Acero, Fernando, et al.
Published: (2024)
by: Acero, Fernando, et al.
Published: (2024)
Deep Reinforcement Learning Agents are not even close to Human Intelligence
by: Delfosse, Quentin, et al.
Published: (2025)
by: Delfosse, Quentin, et al.
Published: (2025)
Interpretable end-to-end Neurosymbolic Reinforcement Learning agents
by: Grandien, Nils, et al.
Published: (2024)
by: Grandien, Nils, et al.
Published: (2024)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
by: Wüst, Antonia, et al.
Published: (2024)
by: Wüst, Antonia, et al.
Published: (2024)
How Hard is it to Confuse a World Model?
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
The Confusing Instance Principle for Online Linear Quadratic Control
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)
by: Weltevrede, Max, et al.
Published: (2025)
Interpretable Policy Distillation for Power Grid Topology Control
by: Dmitruka, Aleksandra, et al.
Published: (2026)
by: Dmitruka, Aleksandra, et al.
Published: (2026)
Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies
by: Rietz, Finn, et al.
Published: (2024)
by: Rietz, Finn, et al.
Published: (2024)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
Prism: Policy Reuse via Interpretable Strategy Mapping in Reinforcement Learning
by: Pravetz, Thomas
Published: (2026)
by: Pravetz, Thomas
Published: (2026)
Three Pathways to Neurosymbolic Reinforcement Learning with Interpretable Model and Policy Networks
by: Graf, Peter, et al.
Published: (2024)
by: Graf, Peter, et al.
Published: (2024)
IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
by: Qin, Yihao, et al.
Published: (2026)
by: Qin, Yihao, et al.
Published: (2026)
Augmented Bayesian Policy Search
by: Kallel, Mahdi, et al.
Published: (2024)
by: Kallel, Mahdi, et al.
Published: (2024)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
by: Lanier, Michael, et al.
Published: (2024)
by: Lanier, Michael, et al.
Published: (2024)
Evaluation-Time Policy Switching for Offline Reinforcement Learning
by: Neggatu, Natinael Solomon, et al.
Published: (2025)
by: Neggatu, Natinael Solomon, et al.
Published: (2025)
From Explainability to Interpretability: Interpretable Policies in Reinforcement Learning Via Model Explanation
by: Li, Peilang, et al.
Published: (2025)
by: Li, Peilang, et al.
Published: (2025)
Exploiting Hybrid Policy in Reinforcement Learning for Interpretable Temporal Logic Manipulation
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
by: Kobanda, Anthony, et al.
Published: (2025)
by: Kobanda, Anthony, et al.
Published: (2025)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
by: Madhow, Sunil, et al.
Published: (2023)
by: Madhow, Sunil, et al.
Published: (2023)
Dataset Distillation for Offline Reinforcement Learning
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
Continual Policy Distillation of Reinforcement Learning-based Controllers for Soft Robotic In-Hand Manipulation
by: Li, Lanpei, et al.
Published: (2024)
by: Li, Lanpei, et al.
Published: (2024)
Better Decisions through the Right Causal World Model
by: Dillies, Elisabeth, et al.
Published: (2025)
by: Dillies, Elisabeth, et al.
Published: (2025)
Similar Items
-
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024) -
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
by: Kohler, Hector, et al.
Published: (2023) -
Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop
by: Kohler, Hector, et al.
Published: (2024) -
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
by: Kohler, Hector, et al.
Published: (2023) -
When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning
by: Berthelot, Yann, et al.
Published: (2026)