Sub-token ViT Embedding via Stochastic Resonance Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lao, Dong, Wu, Yangchao, Liu, Tian Yu, Wong, Alex, Soatto, Stefano
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917658886144000
author Lao, Dong
Wu, Yangchao
Liu, Tian Yu
Wong, Alex
Soatto, Stefano
author_facet Lao, Dong
Wu, Yangchao
Liu, Tian Yu
Wong, Alex
Soatto, Stefano
contents Vision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This representation trades spatial granularity for embedding dimensionality, and results in semantically rich but spatially coarsely quantized feature maps. In order to retrieve spatial details beneficial to fine-grained inference tasks we propose a training-free method inspired by "stochastic resonance". Specifically, we perform sub-token spatial transformations to the input data, and aggregate the resulting ViT features after applying the inverse transformation. The resulting "Stochastic Resonance Transformer" (SRT) retains the rich semantic information of the original representation, but grounds it on a finer-scale spatial domain, partly mitigating the coarse effect of spatial tokenization. SRT is applicable across any layer of any ViT architecture, consistently boosting performance on several tasks including segmentation, classification, depth estimation, and others by up to 14.9% without the need for any fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2310_03967
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sub-token ViT Embedding via Stochastic Resonance Transformers
Lao, Dong
Wu, Yangchao
Liu, Tian Yu
Wong, Alex
Soatto, Stefano
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This representation trades spatial granularity for embedding dimensionality, and results in semantically rich but spatially coarsely quantized feature maps. In order to retrieve spatial details beneficial to fine-grained inference tasks we propose a training-free method inspired by "stochastic resonance". Specifically, we perform sub-token spatial transformations to the input data, and aggregate the resulting ViT features after applying the inverse transformation. The resulting "Stochastic Resonance Transformer" (SRT) retains the rich semantic information of the original representation, but grounds it on a finer-scale spatial domain, partly mitigating the coarse effect of spatial tokenization. SRT is applicable across any layer of any ViT architecture, consistently boosting performance on several tasks including segmentation, classification, depth estimation, and others by up to 14.9% without the need for any fine-tuning.
title Sub-token ViT Embedding via Stochastic Resonance Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2310.03967