AgileIR: Memory-Efficient Group Shifted Windows Attention for Agile Image Restoration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cai, Hongyi, Rahman, Mohammad Mahdinur, Akhtar, Mohammad Shahid, Li, Jie, Wu, Jingyu, Fang, Zhili
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914946013462528
author Cai, Hongyi
Rahman, Mohammad Mahdinur
Akhtar, Mohammad Shahid
Li, Jie
Wu, Jingyu
Fang, Zhili
author_facet Cai, Hongyi
Rahman, Mohammad Mahdinur
Akhtar, Mohammad Shahid
Li, Jie
Wu, Jingyu
Fang, Zhili
contents Image Transformers show a magnificent success in Image Restoration tasks. Nevertheless, most of transformer-based models are strictly bounded by exorbitant memory occupancy. Our goal is to reduce the memory consumption of Swin Transformer and at the same time speed up the model during training process. Thus, we introduce AgileIR, group shifted attention mechanism along with window attention, which sparsely simplifies the model in architecture. We propose Group Shifted Window Attention (GSWA) to decompose Shift Window Multi-head Self Attention (SW-MSA) and Window Multi-head Self Attention (W-MSA) into groups across their attention heads, contributing to shrinking memory usage in back propagation. In addition to that, we keep shifted window masking and its shifted learnable biases during training, in order to induce the model interacting across windows within the channel. We also re-allocate projection parameters to accelerate attention matrix calculation, which we found a negligible decrease in performance. As a result of experiment, compared with our baseline SwinIR and other efficient quantization models, AgileIR keeps the performance still at 32.20 dB on Set5 evaluation dataset, exceeding other methods with tailor-made efficient methods and saves over 50% memory while a large batch size is employed.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06206
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AgileIR: Memory-Efficient Group Shifted Windows Attention for Agile Image Restoration
Cai, Hongyi
Rahman, Mohammad Mahdinur
Akhtar, Mohammad Shahid
Li, Jie
Wu, Jingyu
Fang, Zhili
Computer Vision and Pattern Recognition
Image Transformers show a magnificent success in Image Restoration tasks. Nevertheless, most of transformer-based models are strictly bounded by exorbitant memory occupancy. Our goal is to reduce the memory consumption of Swin Transformer and at the same time speed up the model during training process. Thus, we introduce AgileIR, group shifted attention mechanism along with window attention, which sparsely simplifies the model in architecture. We propose Group Shifted Window Attention (GSWA) to decompose Shift Window Multi-head Self Attention (SW-MSA) and Window Multi-head Self Attention (W-MSA) into groups across their attention heads, contributing to shrinking memory usage in back propagation. In addition to that, we keep shifted window masking and its shifted learnable biases during training, in order to induce the model interacting across windows within the channel. We also re-allocate projection parameters to accelerate attention matrix calculation, which we found a negligible decrease in performance. As a result of experiment, compared with our baseline SwinIR and other efficient quantization models, AgileIR keeps the performance still at 32.20 dB on Set5 evaluation dataset, exceeding other methods with tailor-made efficient methods and saves over 50% memory while a large batch size is employed.
title AgileIR: Memory-Efficient Group Shifted Windows Attention for Agile Image Restoration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.06206