MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Chengxing, Zhang, Xiaoming, Li, Linze, Fu, Yuqian, Gong, Biao, Li, Tianrui, Zhang, Kai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916656448536576
author Xie, Chengxing
Zhang, Xiaoming
Li, Linze
Fu, Yuqian
Gong, Biao
Li, Tianrui
Zhang, Kai
author_facet Xie, Chengxing
Zhang, Xiaoming
Li, Linze
Fu, Yuqian
Gong, Biao
Li, Tianrui
Zhang, Kai
contents Image super-resolution (SR) has significantly advanced through the adoption of Transformer architectures. However, conventional techniques aimed at enlarging the self-attention window to capture broader contexts come with inherent drawbacks, especially the significantly increased computational demands. Moreover, the feature perception within a fixed-size window of existing models restricts the effective receptive field (ERF) and the intermediate feature diversity. We demonstrate that a flexible integration of attention across diverse spatial extents can yield significant performance enhancements. In line with this insight, we introduce Multi-Range Attention Transformer (MAT) for SR tasks. MAT leverages the computational advantages inherent in dilation operation, in conjunction with self-attention mechanism, to facilitate both multi-range attention (MA) and sparse multi-range attention (SMA), enabling efficient capture of both regional and sparse global features. Combined with local feature extraction, MAT adeptly capture dependencies across various spatial ranges, improving the diversity and efficacy of its feature representations. We also introduce the MSConvStar module, which augments the model's ability for multi-range representation learning. Comprehensive experiments show that our MAT exhibits superior performance to existing state-of-the-art SR models with remarkable efficiency (~3.3 faster than SRFormer-light).
format Preprint
id arxiv_https___arxiv_org_abs_2411_17214
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
Xie, Chengxing
Zhang, Xiaoming
Li, Linze
Fu, Yuqian
Gong, Biao
Li, Tianrui
Zhang, Kai
Computer Vision and Pattern Recognition
Image super-resolution (SR) has significantly advanced through the adoption of Transformer architectures. However, conventional techniques aimed at enlarging the self-attention window to capture broader contexts come with inherent drawbacks, especially the significantly increased computational demands. Moreover, the feature perception within a fixed-size window of existing models restricts the effective receptive field (ERF) and the intermediate feature diversity. We demonstrate that a flexible integration of attention across diverse spatial extents can yield significant performance enhancements. In line with this insight, we introduce Multi-Range Attention Transformer (MAT) for SR tasks. MAT leverages the computational advantages inherent in dilation operation, in conjunction with self-attention mechanism, to facilitate both multi-range attention (MA) and sparse multi-range attention (SMA), enabling efficient capture of both regional and sparse global features. Combined with local feature extraction, MAT adeptly capture dependencies across various spatial ranges, improving the diversity and efficacy of its feature representations. We also introduce the MSConvStar module, which augments the model's ability for multi-range representation learning. Comprehensive experiments show that our MAT exhibits superior performance to existing state-of-the-art SR models with remarkable efficiency (~3.3 faster than SRFormer-light).
title MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.17214