Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
Fuente:
arXiv
Guardado en:
| Autores principales: | Kohler, Hector, Delfosse, Quentin, Radji, Waris, Akrour, Riad, Preux, Philippe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
por: Kohler, Hector, et al.
Publicado: (2024)
por: Kohler, Hector, et al.
Publicado: (2024)
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
por: Kohler, Hector, et al.
Publicado: (2023)
por: Kohler, Hector, et al.
Publicado: (2023)
Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop
por: Kohler, Hector, et al.
Publicado: (2024)
por: Kohler, Hector, et al.
Publicado: (2024)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
por: Kohler, Hector, et al.
Publicado: (2023)
por: Kohler, Hector, et al.
Publicado: (2023)
When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning
por: Berthelot, Yann, et al.
Publicado: (2026)
por: Berthelot, Yann, et al.
Publicado: (2026)
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
por: Driss, Brahim, et al.
Publicado: (2025)
por: Driss, Brahim, et al.
Publicado: (2025)
Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces
por: Kobanda, Anthony, et al.
Publicado: (2026)
por: Kobanda, Anthony, et al.
Publicado: (2026)
StaQ it! Growing neural networks for Policy Mirror Descent
por: Shilova, Alena, et al.
Publicado: (2025)
por: Shilova, Alena, et al.
Publicado: (2025)
Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX
por: Radji, Waris, et al.
Publicado: (2025)
por: Radji, Waris, et al.
Publicado: (2025)
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning
por: Delfosse, Quentin, et al.
Publicado: (2024)
por: Delfosse, Quentin, et al.
Publicado: (2024)
GRAIL: Autonomous Concept Grounding for Neuro-Symbolic Reinforcement Learning
por: Shindo, Hikaru, et al.
Publicado: (2026)
por: Shindo, Hikaru, et al.
Publicado: (2026)
IDEQ: an improved diffusion model for the TSP
por: Basson, Mickael, et al.
Publicado: (2024)
por: Basson, Mickael, et al.
Publicado: (2024)
BlendRL: A Framework for Merging Symbolic and Neural Policy Learning
por: Shindo, Hikaru, et al.
Publicado: (2024)
por: Shindo, Hikaru, et al.
Publicado: (2024)
Deep Reinforcement Learning via Object-Centric Attention
por: Blüml, Jannis, et al.
Publicado: (2025)
por: Blüml, Jannis, et al.
Publicado: (2025)
Boosting deep Reinforcement Learning using pretraining with Logical Options
por: Ye, Zihan, et al.
Publicado: (2026)
por: Ye, Zihan, et al.
Publicado: (2026)
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
por: Delfosse, Quentin, et al.
Publicado: (2023)
por: Delfosse, Quentin, et al.
Publicado: (2023)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
por: Acero, Fernando, et al.
Publicado: (2024)
por: Acero, Fernando, et al.
Publicado: (2024)
Deep Reinforcement Learning Agents are not even close to Human Intelligence
por: Delfosse, Quentin, et al.
Publicado: (2025)
por: Delfosse, Quentin, et al.
Publicado: (2025)
Interpretable end-to-end Neurosymbolic Reinforcement Learning agents
por: Grandien, Nils, et al.
Publicado: (2024)
por: Grandien, Nils, et al.
Publicado: (2024)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
por: Wüst, Antonia, et al.
Publicado: (2024)
por: Wüst, Antonia, et al.
Publicado: (2024)
How Hard is it to Confuse a World Model?
por: Radji, Waris, et al.
Publicado: (2025)
por: Radji, Waris, et al.
Publicado: (2025)
The Confusing Instance Principle for Online Linear Quadratic Control
por: Radji, Waris, et al.
Publicado: (2025)
por: Radji, Waris, et al.
Publicado: (2025)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
por: Weltevrede, Max, et al.
Publicado: (2025)
por: Weltevrede, Max, et al.
Publicado: (2025)
Interpretable Policy Distillation for Power Grid Topology Control
por: Dmitruka, Aleksandra, et al.
Publicado: (2026)
por: Dmitruka, Aleksandra, et al.
Publicado: (2026)
Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies
por: Rietz, Finn, et al.
Publicado: (2024)
por: Rietz, Finn, et al.
Publicado: (2024)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
por: Xu, Charles, et al.
Publicado: (2024)
por: Xu, Charles, et al.
Publicado: (2024)
Prism: Policy Reuse via Interpretable Strategy Mapping in Reinforcement Learning
por: Pravetz, Thomas
Publicado: (2026)
por: Pravetz, Thomas
Publicado: (2026)
Three Pathways to Neurosymbolic Reinforcement Learning with Interpretable Model and Policy Networks
por: Graf, Peter, et al.
Publicado: (2024)
por: Graf, Peter, et al.
Publicado: (2024)
IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
por: Qin, Yihao, et al.
Publicado: (2026)
por: Qin, Yihao, et al.
Publicado: (2026)
Augmented Bayesian Policy Search
por: Kallel, Mahdi, et al.
Publicado: (2024)
por: Kallel, Mahdi, et al.
Publicado: (2024)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
por: Lanier, Michael, et al.
Publicado: (2024)
por: Lanier, Michael, et al.
Publicado: (2024)
Evaluation-Time Policy Switching for Offline Reinforcement Learning
por: Neggatu, Natinael Solomon, et al.
Publicado: (2025)
por: Neggatu, Natinael Solomon, et al.
Publicado: (2025)
From Explainability to Interpretability: Interpretable Policies in Reinforcement Learning Via Model Explanation
por: Li, Peilang, et al.
Publicado: (2025)
por: Li, Peilang, et al.
Publicado: (2025)
Exploiting Hybrid Policy in Reinforcement Learning for Interpretable Temporal Logic Manipulation
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
por: Kobanda, Anthony, et al.
Publicado: (2025)
por: Kobanda, Anthony, et al.
Publicado: (2025)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
por: Madhow, Sunil, et al.
Publicado: (2023)
por: Madhow, Sunil, et al.
Publicado: (2023)
Dataset Distillation for Offline Reinforcement Learning
por: Light, Jonathan, et al.
Publicado: (2024)
por: Light, Jonathan, et al.
Publicado: (2024)
Reinforcement Learning via Self-Distillation
por: Hübotter, Jonas, et al.
Publicado: (2026)
por: Hübotter, Jonas, et al.
Publicado: (2026)
Continual Policy Distillation of Reinforcement Learning-based Controllers for Soft Robotic In-Hand Manipulation
por: Li, Lanpei, et al.
Publicado: (2024)
por: Li, Lanpei, et al.
Publicado: (2024)
Better Decisions through the Right Causal World Model
por: Dillies, Elisabeth, et al.
Publicado: (2025)
por: Dillies, Elisabeth, et al.
Publicado: (2025)
Ejemplares similares
-
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
por: Kohler, Hector, et al.
Publicado: (2024) -
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
por: Kohler, Hector, et al.
Publicado: (2023) -
Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop
por: Kohler, Hector, et al.
Publicado: (2024) -
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
por: Kohler, Hector, et al.
Publicado: (2023) -
When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning
por: Berthelot, Yann, et al.
Publicado: (2026)