MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Wenzhuo, Zhu, Fei, Ma, Shijie, Liu, Cheng-Lin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917677547651072
author Liu, Wenzhuo
Zhu, Fei
Ma, Shijie
Liu, Cheng-Lin
author_facet Liu, Wenzhuo
Zhu, Fei
Ma, Shijie
Liu, Cheng-Lin
contents Although Vision Transformers (ViTs) have recently advanced computer vision tasks significantly, an important real-world problem was overlooked: adapting to variable input resolutions. Typically, images are resized to a fixed resolution, such as 224x224, for efficiency during training and inference. However, uniform input size conflicts with real-world scenarios where images naturally vary in resolution. Modifying the preset resolution of a model may severely degrade the performance. In this work, we propose to enhance the model adaptability to resolution variation by optimizing the patch embedding. The proposed method, called Multi-Scale Patch Embedding (MSPE), substitutes the standard patch embedding with multiple variable-sized patch kernels and selects the best parameters for different resolutions, eliminating the need to resize the original image. Our method does not require high-cost training or modifications to other parts, making it easy to apply to most ViT models. Experiments in image classification, segmentation, and detection tasks demonstrate the effectiveness of MSPE, yielding superior performance on low-resolution inputs and performing comparably on high-resolution inputs with existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18240
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
Liu, Wenzhuo
Zhu, Fei
Ma, Shijie
Liu, Cheng-Lin
Computer Vision and Pattern Recognition
Although Vision Transformers (ViTs) have recently advanced computer vision tasks significantly, an important real-world problem was overlooked: adapting to variable input resolutions. Typically, images are resized to a fixed resolution, such as 224x224, for efficiency during training and inference. However, uniform input size conflicts with real-world scenarios where images naturally vary in resolution. Modifying the preset resolution of a model may severely degrade the performance. In this work, we propose to enhance the model adaptability to resolution variation by optimizing the patch embedding. The proposed method, called Multi-Scale Patch Embedding (MSPE), substitutes the standard patch embedding with multiple variable-sized patch kernels and selects the best parameters for different resolutions, eliminating the need to resize the original image. Our method does not require high-cost training or modifications to other parts, making it easy to apply to most ViT models. Experiments in image classification, segmentation, and detection tasks demonstrate the effectiveness of MSPE, yielding superior performance on low-resolution inputs and performing comparably on high-resolution inputs with existing methods.
title MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18240