A Survey on Explainable Deep Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Zelei, Yu, Jiahao, Xing, Xinyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912228171579392
author Cheng, Zelei
Yu, Jiahao
Xing, Xinyu
author_facet Cheng, Zelei
Yu, Jiahao
Xing, Xinyu
contents Deep Reinforcement Learning (DRL) has achieved remarkable success in sequential decision-making tasks across diverse domains, yet its reliance on black-box neural architectures hinders interpretability, trust, and deployment in high-stakes applications. Explainable Deep Reinforcement Learning (XRL) addresses these challenges by enhancing transparency through feature-level, state-level, dataset-level, and model-level explanation techniques. This survey provides a comprehensive review of XRL methods, evaluates their qualitative and quantitative assessment frameworks, and explores their role in policy refinement, adversarial robustness, and security. Additionally, we examine the integration of reinforcement learning with Large Language Models (LLMs), particularly through Reinforcement Learning from Human Feedback (RLHF), which optimizes AI alignment with human preferences. We conclude by highlighting open research challenges and future directions to advance the development of interpretable, reliable, and accountable DRL systems.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06869
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Explainable Deep Reinforcement Learning
Cheng, Zelei
Yu, Jiahao
Xing, Xinyu
Machine Learning
Artificial Intelligence
Deep Reinforcement Learning (DRL) has achieved remarkable success in sequential decision-making tasks across diverse domains, yet its reliance on black-box neural architectures hinders interpretability, trust, and deployment in high-stakes applications. Explainable Deep Reinforcement Learning (XRL) addresses these challenges by enhancing transparency through feature-level, state-level, dataset-level, and model-level explanation techniques. This survey provides a comprehensive review of XRL methods, evaluates their qualitative and quantitative assessment frameworks, and explores their role in policy refinement, adversarial robustness, and security. Additionally, we examine the integration of reinforcement learning with Large Language Models (LLMs), particularly through Reinforcement Learning from Human Feedback (RLHF), which optimizes AI alignment with human preferences. We conclude by highlighting open research challenges and future directions to advance the development of interpretable, reliable, and accountable DRL systems.
title A Survey on Explainable Deep Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.06869