Saved in:
Bibliographic Details
Main Authors: Mao, Wendong, Zhao, Mingfan, Guan, Jianfeng, Dong, Qiwei, Wang, Zhongfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.11549
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908468047249408
author Mao, Wendong
Zhao, Mingfan
Guan, Jianfeng
Dong, Qiwei
Wang, Zhongfeng
author_facet Mao, Wendong
Zhao, Mingfan
Guan, Jianfeng
Dong, Qiwei
Wang, Zhongfeng
contents Deformable Attention Transformers (DAT) have shown remarkable performance in computer vision tasks by adaptively focusing on informative image regions. However, their data-dependent sampling mechanism introduces irregular memory access patterns, posing significant challenges for efficient hardware deployment. Existing acceleration methods either incur high hardware overhead or compromise model accuracy. To address these issues, this paper proposes a hardware-friendly optimization framework for DAT. First, a neural architecture search (NAS)-based method with a new slicing strategy is proposed to automatically divide the input feature into uniform patches during the inference process, avoiding memory conflicts without modifying model architecture. The method explores the optimal slice configuration by jointly optimizing hardware cost and inference accuracy. Secondly, an FPGA-based verification system is designed to test the performance of this framework on edge-side hardware. Algorithm experiments on the ImageNet-1K dataset demonstrate that our hardware-friendly framework can maintain have only 0.2% accuracy drop compared to the baseline DAT. Hardware experiments on Xilinx FPGA show the proposed method reduces DRAM access times to 18% compared with existing DAT acceleration methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search
Mao, Wendong
Zhao, Mingfan
Guan, Jianfeng
Dong, Qiwei
Wang, Zhongfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Deformable Attention Transformers (DAT) have shown remarkable performance in computer vision tasks by adaptively focusing on informative image regions. However, their data-dependent sampling mechanism introduces irregular memory access patterns, posing significant challenges for efficient hardware deployment. Existing acceleration methods either incur high hardware overhead or compromise model accuracy. To address these issues, this paper proposes a hardware-friendly optimization framework for DAT. First, a neural architecture search (NAS)-based method with a new slicing strategy is proposed to automatically divide the input feature into uniform patches during the inference process, avoiding memory conflicts without modifying model architecture. The method explores the optimal slice configuration by jointly optimizing hardware cost and inference accuracy. Secondly, an FPGA-based verification system is designed to test the performance of this framework on edge-side hardware. Algorithm experiments on the ImageNet-1K dataset demonstrate that our hardware-friendly framework can maintain have only 0.2% accuracy drop compared to the baseline DAT. Hardware experiments on Xilinx FPGA show the proposed method reduces DRAM access times to 18% compared with existing DAT acceleration methods.
title A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.11549