Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Jianing, Panagopoulos, Anastasios, Jayaraman, Dinesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914811318632448
author Qian, Jianing
Panagopoulos, Anastasios
Jayaraman, Dinesh
author_facet Qian, Jianing
Panagopoulos, Anastasios
Jayaraman, Dinesh
contents Generic re-usable pre-trained image representation encoders have become a standard component of methods for many computer vision tasks. As visual representations for robots however, their utility has been limited, leading to a recent wave of efforts to pre-train robotics-specific image encoders that are better suited to robotic tasks than their generic counterparts. We propose Scene Objects From Transformers, abbreviated as SOFT, a wrapper around pre-trained vision transformer (PVT) models that bridges this gap without any further training. Rather than construct representations out of only the final layer activations, SOFT individuates and locates object-like entities from PVT attentions, and describes them with PVT activations, producing an object-centric embedding. Across standard choices of generic pre-trained vision transformers PVT, we demonstrate in each case that policies trained on SOFT(PVT) far outstrip standard PVT representations for manipulation tasks in simulated and real settings, approaching the state-of-the-art robotics-aware representations. Code, appendix and videos: https://sites.google.com/view/robot-soft/
format Preprint
id arxiv_https___arxiv_org_abs_2405_15916
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
Qian, Jianing
Panagopoulos, Anastasios
Jayaraman, Dinesh
Computer Vision and Pattern Recognition
Robotics
Generic re-usable pre-trained image representation encoders have become a standard component of methods for many computer vision tasks. As visual representations for robots however, their utility has been limited, leading to a recent wave of efforts to pre-train robotics-specific image encoders that are better suited to robotic tasks than their generic counterparts. We propose Scene Objects From Transformers, abbreviated as SOFT, a wrapper around pre-trained vision transformer (PVT) models that bridges this gap without any further training. Rather than construct representations out of only the final layer activations, SOFT individuates and locates object-like entities from PVT attentions, and describes them with PVT activations, producing an object-centric embedding. Across standard choices of generic pre-trained vision transformers PVT, we demonstrate in each case that policies trained on SOFT(PVT) far outstrip standard PVT representations for manipulation tasks in simulated and real settings, approaching the state-of-the-art robotics-aware representations. Code, appendix and videos: https://sites.google.com/view/robot-soft/
title Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2405.15916