Progressive Focused Transformer for Single Image Super-Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Long, Wei, Zhou, Xingyu, Zhang, Leheng, Gu, Shuhang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915213832355840
author Long, Wei
Zhou, Xingyu
Zhang, Leheng
Gu, Shuhang
author_facet Long, Wei
Zhou, Xingyu
Zhang, Leheng
Gu, Shuhang
contents Transformer-based methods have achieved remarkable results in image super-resolution tasks because they can capture non-local dependencies in low-quality input images. However, this feature-intensive modeling approach is computationally expensive because it calculates the similarities between numerous features that are irrelevant to the query features when obtaining attention weights. These unnecessary similarity calculations not only degrade the reconstruction performance but also introduce significant computational overhead. How to accurately identify the features that are important to the current query features and avoid similarity calculations between irrelevant features remains an urgent problem. To address this issue, we propose a novel and effective Progressive Focused Transformer (PFT) that links all isolated attention maps in the network through Progressive Focused Attention (PFA) to focus attention on the most important tokens. PFA not only enables the network to capture more critical similar features, but also significantly reduces the computational cost of the overall network by filtering out irrelevant features before calculating similarities. Extensive experiments demonstrate the effectiveness of the proposed method, achieving state-of-the-art performance on various single image super-resolution benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20337
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Progressive Focused Transformer for Single Image Super-Resolution
Long, Wei
Zhou, Xingyu
Zhang, Leheng
Gu, Shuhang
Computer Vision and Pattern Recognition
Transformer-based methods have achieved remarkable results in image super-resolution tasks because they can capture non-local dependencies in low-quality input images. However, this feature-intensive modeling approach is computationally expensive because it calculates the similarities between numerous features that are irrelevant to the query features when obtaining attention weights. These unnecessary similarity calculations not only degrade the reconstruction performance but also introduce significant computational overhead. How to accurately identify the features that are important to the current query features and avoid similarity calculations between irrelevant features remains an urgent problem. To address this issue, we propose a novel and effective Progressive Focused Transformer (PFT) that links all isolated attention maps in the network through Progressive Focused Attention (PFA) to focus attention on the most important tokens. PFA not only enables the network to capture more critical similar features, but also significantly reduces the computational cost of the overall network by filtering out irrelevant features before calculating similarities. Extensive experiments demonstrate the effectiveness of the proposed method, achieving state-of-the-art performance on various single image super-resolution benchmarks.
title Progressive Focused Transformer for Single Image Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.20337