HSViT: Horizontally Scalable Vision Transformer

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Chenhao, Li, Chang-Tsun, Lim, Chee Peng, Creighton, Douglas
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916325210718208
author Xu, Chenhao
Li, Chang-Tsun
Lim, Chee Peng
Creighton, Douglas
author_facet Xu, Chenhao
Li, Chang-Tsun
Lim, Chee Peng
Creighton, Douglas
contents Due to its deficiency in prior knowledge (inductive bias), Vision Transformer (ViT) requires pre-training on large-scale datasets to perform well. Moreover, the growing layers and parameters in ViT models impede their applicability to devices with limited computing resources. To mitigate the aforementioned challenges, this paper introduces a novel horizontally scalable vision transformer (HSViT) scheme. Specifically, a novel image-level feature embedding is introduced to ViT, where the preserved inductive bias allows the model to eliminate the need for pre-training while outperforming on small datasets. Besides, a novel horizontally scalable architecture is designed, facilitating collaborative model training and inference across multiple computing devices. The experimental results depict that, without pre-training, HSViT achieves up to 10% higher top-1 accuracy than state-of-the-art schemes on small datasets, while providing existing CNN backbones up to 3.1% improvement in top-1 accuracy on ImageNet. The code is available at https://github.com/xuchenhao001/HSViT.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05196
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HSViT: Horizontally Scalable Vision Transformer
Xu, Chenhao
Li, Chang-Tsun
Lim, Chee Peng
Creighton, Douglas
Computer Vision and Pattern Recognition
Due to its deficiency in prior knowledge (inductive bias), Vision Transformer (ViT) requires pre-training on large-scale datasets to perform well. Moreover, the growing layers and parameters in ViT models impede their applicability to devices with limited computing resources. To mitigate the aforementioned challenges, this paper introduces a novel horizontally scalable vision transformer (HSViT) scheme. Specifically, a novel image-level feature embedding is introduced to ViT, where the preserved inductive bias allows the model to eliminate the need for pre-training while outperforming on small datasets. Besides, a novel horizontally scalable architecture is designed, facilitating collaborative model training and inference across multiple computing devices. The experimental results depict that, without pre-training, HSViT achieves up to 10% higher top-1 accuracy than state-of-the-art schemes on small datasets, while providing existing CNN backbones up to 3.1% improvement in top-1 accuracy on ImageNet. The code is available at https://github.com/xuchenhao001/HSViT.
title HSViT: Horizontally Scalable Vision Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.05196