Unsupervised Salient Patch Selection for Data-Efficient Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Zhaohui, Weng, Paul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929234636701696
author Jiang, Zhaohui
Weng, Paul
author_facet Jiang, Zhaohui
Weng, Paul
contents To improve the sample efficiency of vision-based deep reinforcement learning (RL), we propose a novel method, called SPIRL, to automatically extract important patches from input images. Following Masked Auto-Encoders, SPIRL is based on Vision Transformer models pre-trained in a self-supervised fashion to reconstruct images from randomly-sampled patches. These pre-trained models can then be exploited to detect and select salient patches, defined as hard to reconstruct from neighboring patches. In RL, the SPIRL agent processes selected salient patches via an attention module. We empirically validate SPIRL on Atari games to test its data-efficiency against relevant state-of-the-art methods, including some traditional model-based methods and keypoint-based models. In addition, we analyze our model's interpretability capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03329
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unsupervised Salient Patch Selection for Data-Efficient Reinforcement Learning
Jiang, Zhaohui
Weng, Paul
Computer Vision and Pattern Recognition
Artificial Intelligence
To improve the sample efficiency of vision-based deep reinforcement learning (RL), we propose a novel method, called SPIRL, to automatically extract important patches from input images. Following Masked Auto-Encoders, SPIRL is based on Vision Transformer models pre-trained in a self-supervised fashion to reconstruct images from randomly-sampled patches. These pre-trained models can then be exploited to detect and select salient patches, defined as hard to reconstruct from neighboring patches. In RL, the SPIRL agent processes selected salient patches via an attention module. We empirically validate SPIRL on Atari games to test its data-efficiency against relevant state-of-the-art methods, including some traditional model-based methods and keypoint-based models. In addition, we analyze our model's interpretability capabilities.
title Unsupervised Salient Patch Selection for Data-Efficient Reinforcement Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2402.03329