Reinforcement Learning via Auxiliary Task Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Harish, Abhinav Narayan, Heck, Larry, Hanna, Josiah P., Kira, Zsolt, Szot, Andrew
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913403354742784
author Harish, Abhinav Narayan
Heck, Larry
Hanna, Josiah P.
Kira, Zsolt
Szot, Andrew
author_facet Harish, Abhinav Narayan
Heck, Larry
Hanna, Josiah P.
Kira, Zsolt
Szot, Andrew
contents We present Reinforcement Learning via Auxiliary Task Distillation (AuxDistill), a new method that enables reinforcement learning (RL) to perform long-horizon robot control problems by distilling behaviors from auxiliary RL tasks. AuxDistill achieves this by concurrently carrying out multi-task RL with auxiliary tasks, which are easier to learn and relevant to the main task. A weighted distillation loss transfers behaviors from these auxiliary tasks to solve the main task. We demonstrate that AuxDistill can learn a pixels-to-actions policy for a challenging multi-stage embodied object rearrangement task from the environment reward without demonstrations, a learning curriculum, or pre-trained skills. AuxDistill achieves $2.3 \times$ higher success than the previous state-of-the-art baseline in the Habitat Object Rearrangement benchmark and outperforms methods that use pre-trained skills and expert demonstrations.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reinforcement Learning via Auxiliary Task Distillation
Harish, Abhinav Narayan
Heck, Larry
Hanna, Josiah P.
Kira, Zsolt
Szot, Andrew
Machine Learning
Artificial Intelligence
Robotics
We present Reinforcement Learning via Auxiliary Task Distillation (AuxDistill), a new method that enables reinforcement learning (RL) to perform long-horizon robot control problems by distilling behaviors from auxiliary RL tasks. AuxDistill achieves this by concurrently carrying out multi-task RL with auxiliary tasks, which are easier to learn and relevant to the main task. A weighted distillation loss transfers behaviors from these auxiliary tasks to solve the main task. We demonstrate that AuxDistill can learn a pixels-to-actions policy for a challenging multi-stage embodied object rearrangement task from the environment reward without demonstrations, a learning curriculum, or pre-trained skills. AuxDistill achieves $2.3 \times$ higher success than the previous state-of-the-art baseline in the Habitat Object Rearrangement benchmark and outperforms methods that use pre-trained skills and expert demonstrations.
title Reinforcement Learning via Auxiliary Task Distillation
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2406.17168