VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wickramasinghe, Sachini, Parikh, Dhruv, Zhang, Bingyi, Kannan, Rajgopal, Prasanna, Viktor, Busart, Carl
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917631677693952
author Wickramasinghe, Sachini
Parikh, Dhruv
Zhang, Bingyi
Kannan, Rajgopal
Prasanna, Viktor
Busart, Carl
author_facet Wickramasinghe, Sachini
Parikh, Dhruv
Zhang, Bingyi
Kannan, Rajgopal
Prasanna, Viktor
Busart, Carl
contents Synthetic Aperture Radar (SAR) Automatic Target Recognition (ATR) is a key technique used in military applications like remote-sensing image recognition. Vision Transformers (ViTs) are the current state-of-the-art in various computer vision applications, outperforming their CNN counterparts. However, using ViTs for SAR ATR applications is challenging due to (1) standard ViTs require extensive training data to generalize well due to their low locality; the standard SAR datasets, however, have a limited number of labeled training data which reduces the learning capability of ViTs; (2) ViTs have a high parameter count and are computation intensive which makes their deployment on resource-constrained SAR platforms difficult. In this work, we develop a lightweight ViT model that can be trained directly on small datasets without any pre-training by utilizing the Shifted Patch Tokenization (SPT) and Locality Self-Attention (LSA) modules. We directly train this model on SAR datasets which have limited training samples to evaluate its effectiveness for SAR ATR applications. We evaluate our proposed model, that we call VTR (ViT for SAR ATR), on three widely used SAR datasets: MSTAR, SynthWakeSAR, and GBSAR. Further, we propose a novel FPGA accelerator for VTR, in order to enable deployment for real-time SAR ATR applications.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04527
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
Wickramasinghe, Sachini
Parikh, Dhruv
Zhang, Bingyi
Kannan, Rajgopal
Prasanna, Viktor
Busart, Carl
Computer Vision and Pattern Recognition
Artificial Intelligence
Hardware Architecture
Distributed, Parallel, and Cluster Computing
Synthetic Aperture Radar (SAR) Automatic Target Recognition (ATR) is a key technique used in military applications like remote-sensing image recognition. Vision Transformers (ViTs) are the current state-of-the-art in various computer vision applications, outperforming their CNN counterparts. However, using ViTs for SAR ATR applications is challenging due to (1) standard ViTs require extensive training data to generalize well due to their low locality; the standard SAR datasets, however, have a limited number of labeled training data which reduces the learning capability of ViTs; (2) ViTs have a high parameter count and are computation intensive which makes their deployment on resource-constrained SAR platforms difficult. In this work, we develop a lightweight ViT model that can be trained directly on small datasets without any pre-training by utilizing the Shifted Patch Tokenization (SPT) and Locality Self-Attention (LSA) modules. We directly train this model on SAR datasets which have limited training samples to evaluate its effectiveness for SAR ATR applications. We evaluate our proposed model, that we call VTR (ViT for SAR ATR), on three widely used SAR datasets: MSTAR, SynthWakeSAR, and GBSAR. Further, we propose a novel FPGA accelerator for VTR, in order to enable deployment for real-time SAR ATR applications.
title VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Hardware Architecture
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2404.04527