TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Olszewski, Jan, Rymarczyk, Dawid, Wójcik, Piotr, Pach, Mateusz, Zieliński, Bartosz
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916495060107264
author Olszewski, Jan
Rymarczyk, Dawid
Wójcik, Piotr
Pach, Mateusz
Zieliński, Bartosz
author_facet Olszewski, Jan
Rymarczyk, Dawid
Wójcik, Piotr
Pach, Mateusz
Zieliński, Bartosz
contents Active Visual Exploration (AVE) optimizes the utilization of robotic resources in real-world scenarios by sequentially selecting the most informative observations. However, modern methods require a high computational budget due to processing the same observations multiple times through the autoencoder transformers. As a remedy, we introduce a novel approach to AVE called TOken REcycling (TORE). It divides the encoder into extractor and aggregator components. The extractor processes each observation separately, enabling the reuse of tokens passed to the aggregator. Moreover, to further reduce the computations, we decrease the decoder to only one block. Through extensive experiments, we demonstrate that TORE outperforms state-of-the-art methods while reducing computational overhead by up to 90\%.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15335
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
Olszewski, Jan
Rymarczyk, Dawid
Wójcik, Piotr
Pach, Mateusz
Zieliński, Bartosz
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Active Visual Exploration (AVE) optimizes the utilization of robotic resources in real-world scenarios by sequentially selecting the most informative observations. However, modern methods require a high computational budget due to processing the same observations multiple times through the autoencoder transformers. As a remedy, we introduce a novel approach to AVE called TOken REcycling (TORE). It divides the encoder into extractor and aggregator components. The extractor processes each observation separately, enabling the reuse of tokens passed to the aggregator. Moreover, to further reduce the computations, we decrease the decoder to only one block. Through extensive experiments, we demonstrate that TORE outperforms state-of-the-art methods while reducing computational overhead by up to 90\%.
title TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.15335