Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Xue, Schukat, Michael, Lu, Junlin, Mannion, Patrick, Mason, Karl, Howley, Enda
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911329739079680
author Yang, Xue
Schukat, Michael
Lu, Junlin
Mannion, Patrick
Mason, Karl
Howley, Enda
author_facet Yang, Xue
Schukat, Michael
Lu, Junlin
Mannion, Patrick
Mason, Karl
Howley, Enda
contents Reinforcement learning (RL) excels in various applications but struggles in dynamic environments where the underlying Markov decision process evolves. Continual reinforcement learning (CRL) enables RL agents to continually learn and adapt to new tasks, but balancing stability (preserving prior knowledge) and plasticity (acquiring new knowledge) remains challenging. Existing methods primarily address the stability-plasticity dilemma through mechanisms where past knowledge influences optimization but rarely affects the agent's behavior directly, which may hinder effective knowledge reuse and efficient learning. In contrast, we propose demonstration-guided continual reinforcement learning (DGCRL), which stores prior knowledge in an external, self-evolving demonstration repository that directly guides RL exploration and adaptation. For each task, the agent dynamically selects the most relevant demonstration and follows a curriculum-based strategy to accelerate learning, gradually shifting from demonstration-guided exploration to fully self-exploration. Extensive experiments on 2D navigation and MuJoCo locomotion tasks demonstrate its superior average performance, enhanced knowledge transfer, mitigation of forgetting, and training efficiency. The additional sensitivity analysis and ablation study further validate its effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
Yang, Xue
Schukat, Michael
Lu, Junlin
Mannion, Patrick
Mason, Karl
Howley, Enda
Machine Learning
Reinforcement learning (RL) excels in various applications but struggles in dynamic environments where the underlying Markov decision process evolves. Continual reinforcement learning (CRL) enables RL agents to continually learn and adapt to new tasks, but balancing stability (preserving prior knowledge) and plasticity (acquiring new knowledge) remains challenging. Existing methods primarily address the stability-plasticity dilemma through mechanisms where past knowledge influences optimization but rarely affects the agent's behavior directly, which may hinder effective knowledge reuse and efficient learning. In contrast, we propose demonstration-guided continual reinforcement learning (DGCRL), which stores prior knowledge in an external, self-evolving demonstration repository that directly guides RL exploration and adaptation. For each task, the agent dynamically selects the most relevant demonstration and follows a curriculum-based strategy to accelerate learning, gradually shifting from demonstration-guided exploration to fully self-exploration. Extensive experiments on 2D navigation and MuJoCo locomotion tasks demonstrate its superior average performance, enhanced knowledge transfer, mitigation of forgetting, and training efficiency. The additional sensitivity analysis and ablation study further validate its effectiveness.
title Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
topic Machine Learning
url https://arxiv.org/abs/2512.18670