Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jinchen, Zhang, Mingjian, Zheng, Ling, Weng, Shizhuang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913260779864064
author Zhu, Jinchen
Zhang, Mingjian
Zheng, Ling
Weng, Shizhuang
author_facet Zhu, Jinchen
Zhang, Mingjian
Zheng, Ling
Weng, Shizhuang
contents Recently, the methods based on implicit neural representations have shown excellent capabilities for arbitrary-scale super-resolution (ASSR). Although these methods represent the features of an image by generating latent codes, these latent codes are difficult to adapt for different magnification factors of super-resolution, which seriously affects their performance. Addressing this, we design Multi-Scale Implicit Transformer (MSIT), consisting of an Multi-scale Neural Operator (MSNO) and Multi-Scale Self-Attention (MSSA). Among them, MSNO obtains multi-scale latent codes through feature enhancement, multi-scale characteristics extraction, and multi-scale characteristics merging. MSSA further enhances the multi-scale characteristics of latent codes, resulting in better performance. Furthermore, to improve the performance of network, we propose the Re-Interaction Module (RIM) combined with the cumulative training strategy to improve the diversity of learned information for the network. We have systematically introduced multi-scale characteristics for the first time in ASSR, extensive experiments are performed to validate the effectiveness of MSIT, and our method achieves state-of-the-art performance in arbitrary super-resolution tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06536
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution
Zhu, Jinchen
Zhang, Mingjian
Zheng, Ling
Weng, Shizhuang
Computer Vision and Pattern Recognition
Recently, the methods based on implicit neural representations have shown excellent capabilities for arbitrary-scale super-resolution (ASSR). Although these methods represent the features of an image by generating latent codes, these latent codes are difficult to adapt for different magnification factors of super-resolution, which seriously affects their performance. Addressing this, we design Multi-Scale Implicit Transformer (MSIT), consisting of an Multi-scale Neural Operator (MSNO) and Multi-Scale Self-Attention (MSSA). Among them, MSNO obtains multi-scale latent codes through feature enhancement, multi-scale characteristics extraction, and multi-scale characteristics merging. MSSA further enhances the multi-scale characteristics of latent codes, resulting in better performance. Furthermore, to improve the performance of network, we propose the Re-Interaction Module (RIM) combined with the cumulative training strategy to improve the diversity of learned information for the network. We have systematically introduced multi-scale characteristics for the first time in ASSR, extensive experiments are performed to validate the effectiveness of MSIT, and our method achieves state-of-the-art performance in arbitrary super-resolution tasks.
title Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.06536