How Powerful Potential of Attention on Image Restoration?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Cong, Pan, Jinshan, Jin, Yeying, Wang, Liyan, Wang, Wei, Fu, Gang, Ren, Wenqi, Cao, Xiaochun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913266565906432
author Wang, Cong
Pan, Jinshan
Jin, Yeying
Wang, Liyan
Wang, Wei
Fu, Gang
Ren, Wenqi
Cao, Xiaochun
author_facet Wang, Cong
Pan, Jinshan
Jin, Yeying
Wang, Liyan
Wang, Wei
Fu, Gang
Ren, Wenqi
Cao, Xiaochun
contents Transformers have demonstrated their effectiveness in image restoration tasks. Existing Transformer architectures typically comprise two essential components: multi-head self-attention and feed-forward network (FFN). The former captures long-range pixel dependencies, while the latter enables the model to learn complex patterns and relationships in the data. Previous studies have demonstrated that FFNs are key-value memories \cite{geva2020transformer}, which are vital in modern Transformer architectures. In this paper, we conduct an empirical study to explore the potential of attention mechanisms without using FFN and provide novel structures to demonstrate that removing FFN is flexible for image restoration. Specifically, we propose Continuous Scaling Attention (\textbf{CSAttn}), a method that computes attention continuously in three stages without using FFN. To achieve competitive performance, we propose a series of key components within the attention. Our designs provide a closer look at the attention mechanism and reveal that some simple operations can significantly affect the model performance. We apply our \textbf{CSAttn} to several image restoration tasks and show that our model can outperform CNN-based and Transformer-based image restoration approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10336
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Powerful Potential of Attention on Image Restoration?
Wang, Cong
Pan, Jinshan
Jin, Yeying
Wang, Liyan
Wang, Wei
Fu, Gang
Ren, Wenqi
Cao, Xiaochun
Computer Vision and Pattern Recognition
Transformers have demonstrated their effectiveness in image restoration tasks. Existing Transformer architectures typically comprise two essential components: multi-head self-attention and feed-forward network (FFN). The former captures long-range pixel dependencies, while the latter enables the model to learn complex patterns and relationships in the data. Previous studies have demonstrated that FFNs are key-value memories \cite{geva2020transformer}, which are vital in modern Transformer architectures. In this paper, we conduct an empirical study to explore the potential of attention mechanisms without using FFN and provide novel structures to demonstrate that removing FFN is flexible for image restoration. Specifically, we propose Continuous Scaling Attention (\textbf{CSAttn}), a method that computes attention continuously in three stages without using FFN. To achieve competitive performance, we propose a series of key components within the attention. Our designs provide a closer look at the attention mechanism and reveal that some simple operations can significantly affect the model performance. We apply our \textbf{CSAttn} to several image restoration tasks and show that our model can outperform CNN-based and Transformer-based image restoration approaches.
title How Powerful Potential of Attention on Image Restoration?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.10336