QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Grace, Jain, Ayush, Hwang, Injune, Sun, Shao-Hua, Lim, Joseph J.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918002559025152
author Zhang, Grace
Jain, Ayush
Hwang, Injune
Sun, Shao-Hua
Lim, Joseph J.
author_facet Zhang, Grace
Jain, Ayush
Hwang, Injune
Sun, Shao-Hua
Lim, Joseph J.
contents Multi-task reinforcement learning (MTRL) aims to learn several tasks simultaneously for better sample efficiency than learning them separately. Traditional methods achieve this by sharing parameters or relabeled data between tasks. In this work, we introduce a new framework for sharing behavioral policies across tasks, which can be used in addition to existing MTRL methods. The key idea is to improve each task's off-policy data collection by employing behaviors from other task policies. Selectively sharing helpful behaviors acquired in one task to collect training data for another task can lead to higher-quality trajectories, leading to more sample-efficient MTRL. Thus, we introduce a simple and principled framework called Q-switch mixture of policies (QMP) that selectively shares behavior between different task policies by using the task's Q-function to evaluate and select useful shareable behaviors. We theoretically analyze how QMP improves the sample efficiency of the underlying RL algorithm. Our experiments show that QMP's behavioral policy sharing provides complementary gains over many popular MTRL algorithms and outperforms alternative ways to share behaviors in various manipulation, locomotion, and navigation environments. Videos are available at https://qmp-mtrl.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2302_00671
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing
Zhang, Grace
Jain, Ayush
Hwang, Injune
Sun, Shao-Hua
Lim, Joseph J.
Machine Learning
Artificial Intelligence
Robotics
Multi-task reinforcement learning (MTRL) aims to learn several tasks simultaneously for better sample efficiency than learning them separately. Traditional methods achieve this by sharing parameters or relabeled data between tasks. In this work, we introduce a new framework for sharing behavioral policies across tasks, which can be used in addition to existing MTRL methods. The key idea is to improve each task's off-policy data collection by employing behaviors from other task policies. Selectively sharing helpful behaviors acquired in one task to collect training data for another task can lead to higher-quality trajectories, leading to more sample-efficient MTRL. Thus, we introduce a simple and principled framework called Q-switch mixture of policies (QMP) that selectively shares behavior between different task policies by using the task's Q-function to evaluate and select useful shareable behaviors. We theoretically analyze how QMP improves the sample efficiency of the underlying RL algorithm. Our experiments show that QMP's behavioral policy sharing provides complementary gains over many popular MTRL algorithms and outperforms alternative ways to share behaviors in various manipulation, locomotion, and navigation environments. Videos are available at https://qmp-mtrl.github.io.
title QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2302.00671