Dense Dynamics-Aware Reward Synthesis: Integrating Prior Experience with Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koprulu, Cevahir, Li, Po-han, Qiu, Tianyu, Zhao, Ruihan, Westenbroek, Tyler, Fridovich-Keil, David, Chinchali, Sandeep, Topcu, Ufuk
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912993810317312
author Koprulu, Cevahir
Li, Po-han
Qiu, Tianyu
Zhao, Ruihan
Westenbroek, Tyler
Fridovich-Keil, David
Chinchali, Sandeep
Topcu, Ufuk
author_facet Koprulu, Cevahir
Li, Po-han
Qiu, Tianyu
Zhao, Ruihan
Westenbroek, Tyler
Fridovich-Keil, David
Chinchali, Sandeep
Topcu, Ufuk
contents Many continuous control problems can be formulated as sparse-reward reinforcement learning (RL) tasks. In principle, online RL methods can automatically explore the state space to solve each new task. However, discovering sequences of actions that lead to a non-zero reward becomes exponentially more difficult as the task horizon increases. Manually shaping rewards can accelerate learning for a fixed task, but it is an arduous process that must be repeated for each new environment. We introduce a systematic reward-shaping framework that distills the information contained in 1) a task-agnostic prior data set and 2) a small number of task-specific expert demonstrations, and then uses these priors to synthesize dense dynamics-aware rewards for the given task. This supervision substantially accelerates learning in our experiments, and we provide analysis demonstrating how the approach can effectively guide online learning agents to faraway goals.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01114
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dense Dynamics-Aware Reward Synthesis: Integrating Prior Experience with Demonstrations
Koprulu, Cevahir
Li, Po-han
Qiu, Tianyu
Zhao, Ruihan
Westenbroek, Tyler
Fridovich-Keil, David
Chinchali, Sandeep
Topcu, Ufuk
Machine Learning
Many continuous control problems can be formulated as sparse-reward reinforcement learning (RL) tasks. In principle, online RL methods can automatically explore the state space to solve each new task. However, discovering sequences of actions that lead to a non-zero reward becomes exponentially more difficult as the task horizon increases. Manually shaping rewards can accelerate learning for a fixed task, but it is an arduous process that must be repeated for each new environment. We introduce a systematic reward-shaping framework that distills the information contained in 1) a task-agnostic prior data set and 2) a small number of task-specific expert demonstrations, and then uses these priors to synthesize dense dynamics-aware rewards for the given task. This supervision substantially accelerates learning in our experiments, and we provide analysis demonstrating how the approach can effectively guide online learning agents to faraway goals.
title Dense Dynamics-Aware Reward Synthesis: Integrating Prior Experience with Demonstrations
topic Machine Learning
url https://arxiv.org/abs/2412.01114