Pre-trained Visual Dynamics Representations for Efficient Policy Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Hao, Zhou, Bohan, Lu, Zongqing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909377977384960
author Luo, Hao
Zhou, Bohan
Lu, Zongqing
author_facet Luo, Hao
Zhou, Bohan
Lu, Zongqing
contents Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence of action annotations and the common domain gap with downstream tasks hinder utilizing videos for RL pre-training. To address the challenge of pre-training with videos, we propose Pre-trained Visual Dynamics Representations (PVDR) to bridge the domain gap between videos and downstream tasks for efficient policy learning. By adopting video prediction as a pre-training task, we use a Transformer-based Conditional Variational Autoencoder (CVAE) to learn visual dynamics representations. The pre-trained visual dynamics representations capture the visual dynamics prior knowledge in the videos. This abstract prior knowledge can be readily adapted to downstream tasks and aligned with executable actions through online adaptation. We conduct experiments on a series of robotics visual control tasks and verify that PVDR is an effective form for pre-training with videos to promote policy learning.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03169
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pre-trained Visual Dynamics Representations for Efficient Policy Learning
Luo, Hao
Zhou, Bohan
Lu, Zongqing
Computer Vision and Pattern Recognition
Machine Learning
Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence of action annotations and the common domain gap with downstream tasks hinder utilizing videos for RL pre-training. To address the challenge of pre-training with videos, we propose Pre-trained Visual Dynamics Representations (PVDR) to bridge the domain gap between videos and downstream tasks for efficient policy learning. By adopting video prediction as a pre-training task, we use a Transformer-based Conditional Variational Autoencoder (CVAE) to learn visual dynamics representations. The pre-trained visual dynamics representations capture the visual dynamics prior knowledge in the videos. This abstract prior knowledge can be readily adapted to downstream tasks and aligned with executable actions through online adaptation. We conduct experiments on a series of robotics visual control tasks and verify that PVDR is an effective form for pre-training with videos to promote policy learning.
title Pre-trained Visual Dynamics Representations for Efficient Policy Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.03169