Exploration by Random Reward Perturbation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Haozhe, Fu, Guoji, Luo, Zhengding, Wu, Jiele, Leong, Tze-Yun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915335381188608
author Ma, Haozhe
Fu, Guoji
Luo, Zhengding
Wu, Jiele
Leong, Tze-Yun
author_facet Ma, Haozhe
Fu, Guoji
Luo, Zhengding
Wu, Jiele
Leong, Tze-Yun
contents We introduce Random Reward Perturbation (RRP), a novel exploration strategy for reinforcement learning (RL). Our theoretical analyses demonstrate that adding zero-mean noise to environmental rewards effectively enhances policy diversity during training, thereby expanding the range of exploration. RRP is fully compatible with the action-perturbation-based exploration strategies, such as $ε$-greedy, stochastic policies, and entropy regularization, providing additive improvements to exploration effects. It is general, lightweight, and can be integrated into existing RL algorithms with minimal implementation effort and negligible computational overhead. RRP establishes a theoretical connection between reward shaping and noise-driven exploration, highlighting their complementary potential. Experiments show that RRP significantly boosts the performance of Proximal Policy Optimization and Soft Actor-Critic, achieving higher sample efficiency and escaping local optima across various tasks, under both sparse and dense reward scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploration by Random Reward Perturbation
Ma, Haozhe
Fu, Guoji
Luo, Zhengding
Wu, Jiele
Leong, Tze-Yun
Machine Learning
Artificial Intelligence
We introduce Random Reward Perturbation (RRP), a novel exploration strategy for reinforcement learning (RL). Our theoretical analyses demonstrate that adding zero-mean noise to environmental rewards effectively enhances policy diversity during training, thereby expanding the range of exploration. RRP is fully compatible with the action-perturbation-based exploration strategies, such as $ε$-greedy, stochastic policies, and entropy regularization, providing additive improvements to exploration effects. It is general, lightweight, and can be integrated into existing RL algorithms with minimal implementation effort and negligible computational overhead. RRP establishes a theoretical connection between reward shaping and noise-driven exploration, highlighting their complementary potential. Experiments show that RRP significantly boosts the performance of Proximal Policy Optimization and Soft Actor-Critic, achieving higher sample efficiency and escaping local optima across various tasks, under both sparse and dense reward scenarios.
title Exploration by Random Reward Perturbation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.08737