ViT-1.58b: Mobile Vision Transformers in the 1-bit Era

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Zhengqing, Zhou, Rong, Wang, Hongyi, He, Lifang, Ye, Yanfang, Sun, Lichao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909231689498624
author Yuan, Zhengqing
Zhou, Rong
Wang, Hongyi
He, Lifang
Ye, Yanfang
Sun, Lichao
author_facet Yuan, Zhengqing
Zhou, Rong
Wang, Hongyi
He, Lifang
Ye, Yanfang
Sun, Lichao
contents Vision Transformers (ViTs) have achieved remarkable performance in various image classification tasks by leveraging the attention mechanism to process image patches as tokens. However, the high computational and memory demands of ViTs pose significant challenges for deployment in resource-constrained environments. This paper introduces ViT-1.58b, a novel 1.58-bit quantized ViT model designed to drastically reduce memory and computational overhead while preserving competitive performance. ViT-1.58b employs ternary quantization, which refines the balance between efficiency and accuracy by constraining weights to {-1, 0, 1} and quantizing activations to 8-bit precision. Our approach ensures efficient scaling in terms of both memory and computation. Experiments on CIFAR-10 and ImageNet-1k demonstrate that ViT-1.58b maintains comparable accuracy to full-precision Vit, with significant reductions in memory usage and computational costs. This paper highlights the potential of extreme quantization techniques in developing sustainable AI solutions and contributes to the broader discourse on efficient model deployment in practical applications. Our code and weights are available at https://github.com/DLYuanGod/ViT-1.58b.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18051
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
Yuan, Zhengqing
Zhou, Rong
Wang, Hongyi
He, Lifang
Ye, Yanfang
Sun, Lichao
Computer Vision and Pattern Recognition
Vision Transformers (ViTs) have achieved remarkable performance in various image classification tasks by leveraging the attention mechanism to process image patches as tokens. However, the high computational and memory demands of ViTs pose significant challenges for deployment in resource-constrained environments. This paper introduces ViT-1.58b, a novel 1.58-bit quantized ViT model designed to drastically reduce memory and computational overhead while preserving competitive performance. ViT-1.58b employs ternary quantization, which refines the balance between efficiency and accuracy by constraining weights to {-1, 0, 1} and quantizing activations to 8-bit precision. Our approach ensures efficient scaling in terms of both memory and computation. Experiments on CIFAR-10 and ImageNet-1k demonstrate that ViT-1.58b maintains comparable accuracy to full-precision Vit, with significant reductions in memory usage and computational costs. This paper highlights the potential of extreme quantization techniques in developing sustainable AI solutions and contributes to the broader discourse on efficient model deployment in practical applications. Our code and weights are available at https://github.com/DLYuanGod/ViT-1.58b.
title ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.18051