Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models
Fuente:
arXiv
Saved in:
| Main Authors: | Deproost, Senne, Steckelmacher, Denis, Nowé, Ann |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning
by: Deproost, Senne, et al.
Published: (2025)
by: Deproost, Senne, et al.
Published: (2025)
Human-Readable Programs as Actors of Reinforcement Learning Agents Using Critic-Moderated Evolution
by: Deproost, Senne, et al.
Published: (2024)
by: Deproost, Senne, et al.
Published: (2024)
Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies
by: Deproost, Senne, et al.
Published: (2026)
by: Deproost, Senne, et al.
Published: (2026)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
Fairness-Aware Reinforcement Learning (FAReL): A Framework for Transparent and Balanced Sequential Decision-Making
by: Cimpean, Alexandra, et al.
Published: (2025)
by: Cimpean, Alexandra, et al.
Published: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
Evaluating COVID-19 vaccine allocation policies using Bayesian $m$-top exploration
by: Cimpean, Alexandra, et al.
Published: (2023)
by: Cimpean, Alexandra, et al.
Published: (2023)
Online Planning in POMDPs with State-Requests
by: Avalos, Raphael, et al.
Published: (2024)
by: Avalos, Raphael, et al.
Published: (2024)
Deep RL With Information Constrained Policies: Generalization in Continuous Control
by: Malloy, Tailia, et al.
Published: (2020)
by: Malloy, Tailia, et al.
Published: (2020)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
APC-RL: Exceeding Data-Driven Behavior Priors with Adaptive Policy Composition
by: Rietz, Finn, et al.
Published: (2026)
by: Rietz, Finn, et al.
Published: (2026)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
by: Nie, Dong
Published: (2026)
by: Nie, Dong
Published: (2026)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
by: Samadi, Amir, et al.
Published: (2024)
by: Samadi, Amir, et al.
Published: (2024)
Preference Guided Iterated Pareto Referent Optimisation for Accessible Route Planning
by: Speziali, Paolo, et al.
Published: (2026)
by: Speziali, Paolo, et al.
Published: (2026)
Laser Learning Environment: A new environment for coordination-critical multi-agent tasks
by: Molinghen, Yannick, et al.
Published: (2024)
by: Molinghen, Yannick, et al.
Published: (2024)
Variational Distillation of Diffusion Policies into Mixture of Experts
by: Zhou, Hongyi, et al.
Published: (2024)
by: Zhou, Hongyi, et al.
Published: (2024)
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
by: Sander, Jacob, et al.
Published: (2026)
by: Sander, Jacob, et al.
Published: (2026)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
Causally-informed Deep Learning towards Explainable and Generalizable Outcomes Prediction in Critical Care
by: Cheng, Yuxiao, et al.
Published: (2025)
by: Cheng, Yuxiao, et al.
Published: (2025)
Can We Optimize Deep RL Policy Weights as Trajectory Modeling?
by: Tang, Hongyao
Published: (2025)
by: Tang, Hongyao
Published: (2025)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
by: Gu, Hao, et al.
Published: (2026)
by: Gu, Hao, et al.
Published: (2026)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Explainable AI Based Diagnosis of Poisoning Attacks in Evolutionary Swarms
by: Asadi, Mehrdad, et al.
Published: (2025)
by: Asadi, Mehrdad, et al.
Published: (2025)
Proximal Policy Distillation
by: Spigler, Giacomo
Published: (2024)
by: Spigler, Giacomo
Published: (2024)
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
by: Gross, Dennis, et al.
Published: (2025)
by: Gross, Dennis, et al.
Published: (2025)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
by: Yang, Zhuolin, et al.
Published: (2026)
by: Yang, Zhuolin, et al.
Published: (2026)
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
by: Sen, Sayambhu, et al.
Published: (2025)
by: Sen, Sayambhu, et al.
Published: (2025)
Scalable Policy-Based RL Algorithms for POMDPs
by: Anjarlekar, Ameya, et al.
Published: (2025)
by: Anjarlekar, Ameya, et al.
Published: (2025)
The Impact of Quantization and Pruning on Deep Reinforcement Learning Models
by: Lu, Heng, et al.
Published: (2024)
by: Lu, Heng, et al.
Published: (2024)
Feasibility-Aware Decision-Focused Learning for Predicting Parameters in the Constraints
by: Mandi, Jayanta, et al.
Published: (2025)
by: Mandi, Jayanta, et al.
Published: (2025)
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers
by: Guiducci, Leonardo, et al.
Published: (2025)
by: Guiducci, Leonardo, et al.
Published: (2025)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
by: Humayoo, Mahammad, et al.
Published: (2018)
by: Humayoo, Mahammad, et al.
Published: (2018)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Causal State Distillation for Explainable Reinforcement Learning
by: Lu, Wenhao, et al.
Published: (2023)
by: Lu, Wenhao, et al.
Published: (2023)
ShuttleEnv: An Interactive Data-Driven RL Environment for Badminton Strategy Modeling
by: Li, Ang, et al.
Published: (2026)
by: Li, Ang, et al.
Published: (2026)
Similar Items
-
Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning
by: Deproost, Senne, et al.
Published: (2025) -
Human-Readable Programs as Actors of Reinforcement Learning Agents Using Critic-Moderated Evolution
by: Deproost, Senne, et al.
Published: (2024) -
Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies
by: Deproost, Senne, et al.
Published: (2026) -
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024) -
Fairness-Aware Reinforcement Learning (FAReL): A Framework for Transparent and Balanced Sequential Decision-Making
by: Cimpean, Alexandra, et al.
Published: (2025)