Automatic Environment Shaping is the Next Frontier in RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Younghyo, Margolis, Gabriel B., Agrawal, Pulkit
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911964866805760
author Park, Younghyo
Margolis, Gabriel B.
Agrawal, Pulkit
author_facet Park, Younghyo
Margolis, Gabriel B.
Agrawal, Pulkit
contents Many roboticists dream of presenting a robot with a task in the evening and returning the next morning to find the robot capable of solving the task. What is preventing us from achieving this? Sim-to-real reinforcement learning (RL) has achieved impressive performance on challenging robotics tasks, but requires substantial human effort to set up the task in a way that is amenable to RL. It's our position that algorithmic improvements in policy optimization and other ideas should be guided towards resolving the primary bottleneck of shaping the training environment, i.e., designing observations, actions, rewards and simulation dynamics. Most practitioners don't tune the RL algorithm, but other environment parameters to obtain a desirable controller. We posit that scaling RL to diverse robotic tasks will only be achieved if the community focuses on automating environment shaping procedures.
format Preprint
id arxiv_https___arxiv_org_abs_2407_16186
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automatic Environment Shaping is the Next Frontier in RL
Park, Younghyo
Margolis, Gabriel B.
Agrawal, Pulkit
Robotics
Artificial Intelligence
Machine Learning
Many roboticists dream of presenting a robot with a task in the evening and returning the next morning to find the robot capable of solving the task. What is preventing us from achieving this? Sim-to-real reinforcement learning (RL) has achieved impressive performance on challenging robotics tasks, but requires substantial human effort to set up the task in a way that is amenable to RL. It's our position that algorithmic improvements in policy optimization and other ideas should be guided towards resolving the primary bottleneck of shaping the training environment, i.e., designing observations, actions, rewards and simulation dynamics. Most practitioners don't tune the RL algorithm, but other environment parameters to obtain a desirable controller. We posit that scaling RL to diverse robotic tasks will only be achieved if the community focuses on automating environment shaping procedures.
title Automatic Environment Shaping is the Next Frontier in RL
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.16186