ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chen Bo Calvin, Hong, Zhang-Wei, Pacchiano, Aldo, Agrawal, Pulkit
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915170862759936
author Zhang, Chen Bo Calvin
Hong, Zhang-Wei
Pacchiano, Aldo
Agrawal, Pulkit
author_facet Zhang, Chen Bo Calvin
Hong, Zhang-Wei
Pacchiano, Aldo
Agrawal, Pulkit
contents Reward shaping is critical in reinforcement learning (RL), particularly for complex tasks where sparse rewards can hinder learning. However, choosing effective shaping rewards from a set of reward functions in a computationally efficient manner remains an open challenge. We propose Online Reward Selection and Policy Optimization (ORSO), a novel approach that frames the selection of shaping reward function as an online model selection problem. ORSO automatically identifies performant shaping reward functions without human intervention with provable regret guarantees. We demonstrate ORSO's effectiveness across various continuous control tasks. Compared to prior approaches, ORSO significantly reduces the amount of data required to evaluate a shaping reward function, resulting in superior data efficiency and a significant reduction in computational time (up to 8 times). ORSO consistently identifies high-quality reward functions outperforming prior methods by more than 50% and on average identifies policies as performant as the ones learned using manually engineered reward functions by domain experts.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13837
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
Zhang, Chen Bo Calvin
Hong, Zhang-Wei
Pacchiano, Aldo
Agrawal, Pulkit
Machine Learning
Artificial Intelligence
Robotics
Reward shaping is critical in reinforcement learning (RL), particularly for complex tasks where sparse rewards can hinder learning. However, choosing effective shaping rewards from a set of reward functions in a computationally efficient manner remains an open challenge. We propose Online Reward Selection and Policy Optimization (ORSO), a novel approach that frames the selection of shaping reward function as an online model selection problem. ORSO automatically identifies performant shaping reward functions without human intervention with provable regret guarantees. We demonstrate ORSO's effectiveness across various continuous control tasks. Compared to prior approaches, ORSO significantly reduces the amount of data required to evaluate a shaping reward function, resulting in superior data efficiency and a significant reduction in computational time (up to 8 times). ORSO consistently identifies high-quality reward functions outperforming prior methods by more than 50% and on average identifies policies as performant as the ones learned using manually engineered reward functions by domain experts.
title ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2410.13837