Saved in:
Bibliographic Details
Main Authors: Nguyen, Duy-Kien, Assran, Mahmoud, Jain, Unnat, Oswald, Martin R., Snoek, Cees G. M., Chen, Xinlei
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.09415
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916652042420224
author Nguyen, Duy-Kien
Assran, Mahmoud
Jain, Unnat
Oswald, Martin R.
Snoek, Cees G. M.
Chen, Xinlei
author_facet Nguyen, Duy-Kien
Assran, Mahmoud
Jain, Unnat
Oswald, Martin R.
Snoek, Cees G. M.
Chen, Xinlei
contents This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias of locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a token and achieve highly performant results. This is substantially different from the popular design in Vision Transformer, which maintains the inductive bias from ConvNets towards local neighborhoods (e.g. by treating each 16x16 patch as a token). We showcase the effectiveness of pixels-as-tokens across three well-studied computer vision tasks: supervised learning for classification and regression, self-supervised learning via masked autoencoding, and image generation with diffusion models. Although it's computationally less practical to directly operate on individual pixels, we believe the community must be made aware of this surprising piece of knowledge when devising the next generation of neural network architectures for computer vision.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09415
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
Nguyen, Duy-Kien
Assran, Mahmoud
Jain, Unnat
Oswald, Martin R.
Snoek, Cees G. M.
Chen, Xinlei
Computer Vision and Pattern Recognition
Machine Learning
This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias of locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a token and achieve highly performant results. This is substantially different from the popular design in Vision Transformer, which maintains the inductive bias from ConvNets towards local neighborhoods (e.g. by treating each 16x16 patch as a token). We showcase the effectiveness of pixels-as-tokens across three well-studied computer vision tasks: supervised learning for classification and regression, self-supervised learning via masked autoencoding, and image generation with diffusion models. Although it's computationally less practical to directly operate on individual pixels, we believe the community must be made aware of this surprising piece of knowledge when devising the next generation of neural network architectures for computer vision.
title An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.09415