FasterViT: Fast Vision Transformers with Hierarchical Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hatamizadeh, Ali, Heinrich, Greg, Yin, Hongxu, Tao, Andrew, Alvarez, Jose M., Kautz, Jan, Molchanov, Pavlo
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917627795865600
author Hatamizadeh, Ali
Heinrich, Greg
Yin, Hongxu
Tao, Andrew
Alvarez, Jose M.
Kautz, Jan
Molchanov, Pavlo
author_facet Hatamizadeh, Ali
Heinrich, Greg
Yin, Hongxu
Tao, Andrew
Alvarez, Jose M.
Kautz, Jan
Molchanov, Pavlo
contents We design a new family of hybrid CNN-ViT neural networks, named FasterViT, with a focus on high image throughput for computer vision (CV) applications. FasterViT combines the benefits of fast local representation learning in CNNs and global modeling properties in ViT. Our newly introduced Hierarchical Attention (HAT) approach decomposes global self-attention with quadratic complexity into a multi-level attention with reduced computational costs. We benefit from efficient window-based self-attention. Each window has access to dedicated carrier tokens that participate in local and global representation learning. At a high level, global self-attentions enable the efficient cross-window communication at lower costs. FasterViT achieves a SOTA Pareto-front in terms of accuracy and image throughput. We have extensively validated its effectiveness on various CV tasks including classification, object detection and segmentation. We also show that HAT can be used as a plug-and-play module for existing networks and enhance them. We further demonstrate significantly faster and more accurate performance than competitive counterparts for images with high resolution. Code is available at https://github.com/NVlabs/FasterViT.
format Preprint
id arxiv_https___arxiv_org_abs_2306_06189
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FasterViT: Fast Vision Transformers with Hierarchical Attention
Hatamizadeh, Ali
Heinrich, Greg
Yin, Hongxu
Tao, Andrew
Alvarez, Jose M.
Kautz, Jan
Molchanov, Pavlo
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We design a new family of hybrid CNN-ViT neural networks, named FasterViT, with a focus on high image throughput for computer vision (CV) applications. FasterViT combines the benefits of fast local representation learning in CNNs and global modeling properties in ViT. Our newly introduced Hierarchical Attention (HAT) approach decomposes global self-attention with quadratic complexity into a multi-level attention with reduced computational costs. We benefit from efficient window-based self-attention. Each window has access to dedicated carrier tokens that participate in local and global representation learning. At a high level, global self-attentions enable the efficient cross-window communication at lower costs. FasterViT achieves a SOTA Pareto-front in terms of accuracy and image throughput. We have extensively validated its effectiveness on various CV tasks including classification, object detection and segmentation. We also show that HAT can be used as a plug-and-play module for existing networks and enhance them. We further demonstrate significantly faster and more accurate performance than competitive counterparts for images with high resolution. Code is available at https://github.com/NVlabs/FasterViT.
title FasterViT: Fast Vision Transformers with Hierarchical Attention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2306.06189