HAAT: Hybrid Attention Aggregation Transformer for Image Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lai, Song-Jiang, Cheung, Tsun-Hin, Fung, Ka-Chun, Xue, Kai-wen, Lam, Kin-Man
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929620979286016
author Lai, Song-Jiang
Cheung, Tsun-Hin
Fung, Ka-Chun
Xue, Kai-wen
Lam, Kin-Man
author_facet Lai, Song-Jiang
Cheung, Tsun-Hin
Fung, Ka-Chun
Xue, Kai-wen
Lam, Kin-Man
contents In the research area of image super-resolution, Swin-transformer-based models are favored for their global spatial modeling and shifting window attention mechanism. However, existing methods often limit self-attention to non overlapping windows to cut costs and ignore the useful information that exists across channels. To address this issue, this paper introduces a novel model, the Hybrid Attention Aggregation Transformer (HAAT), designed to better leverage feature information. HAAT is constructed by integrating Swin-Dense-Residual-Connected Blocks (SDRCB) with Hybrid Grid Attention Blocks (HGAB). SDRCB expands the receptive field while maintaining a streamlined architecture, resulting in enhanced performance. HGAB incorporates channel attention, sparse attention, and window attention to improve nonlocal feature fusion and achieve more visually compelling results. Experimental evaluations demonstrate that HAAT surpasses state-of-the-art methods on benchmark datasets. Keywords: Image super-resolution, Computer vision, Attention mechanism, Transformer
format Preprint
id arxiv_https___arxiv_org_abs_2411_18003
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HAAT: Hybrid Attention Aggregation Transformer for Image Super-Resolution
Lai, Song-Jiang
Cheung, Tsun-Hin
Fung, Ka-Chun
Xue, Kai-wen
Lam, Kin-Man
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
In the research area of image super-resolution, Swin-transformer-based models are favored for their global spatial modeling and shifting window attention mechanism. However, existing methods often limit self-attention to non overlapping windows to cut costs and ignore the useful information that exists across channels. To address this issue, this paper introduces a novel model, the Hybrid Attention Aggregation Transformer (HAAT), designed to better leverage feature information. HAAT is constructed by integrating Swin-Dense-Residual-Connected Blocks (SDRCB) with Hybrid Grid Attention Blocks (HGAB). SDRCB expands the receptive field while maintaining a streamlined architecture, resulting in enhanced performance. HGAB incorporates channel attention, sparse attention, and window attention to improve nonlocal feature fusion and achieve more visually compelling results. Experimental evaluations demonstrate that HAAT surpasses state-of-the-art methods on benchmark datasets. Keywords: Image super-resolution, Computer vision, Attention mechanism, Transformer
title HAAT: Hybrid Attention Aggregation Transformer for Image Super-Resolution
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.18003