Xformer: Hybrid X-Shaped Transformer for Image Denoising

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jiale, Zhang, Yulun, Gu, Jinjin, Dong, Jiahua, Kong, Linghe, Yang, Xiaokang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914690579300352
author Zhang, Jiale
Zhang, Yulun
Gu, Jinjin
Dong, Jiahua
Kong, Linghe
Yang, Xiaokang
author_facet Zhang, Jiale
Zhang, Yulun
Gu, Jinjin
Dong, Jiahua
Kong, Linghe
Yang, Xiaokang
contents In this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two types of Transformer blocks. The spatial-wise Transformer block performs fine-grained local patches interactions across tokens defined by spatial dimension. The channel-wise Transformer block performs direct global context interactions across tokens defined by channel dimension. Based on the concurrent network structure, we design two branches to conduct these two interaction fashions. Within each branch, we employ an encoder-decoder architecture to capture multi-scale features. Besides, we propose the Bidirectional Connection Unit (BCU) to couple the learned representations from these two branches while providing enhanced information fusion. The joint designs make our Xformer powerful to conduct global information modeling in both spatial and channel dimensions. Extensive experiments show that Xformer, under the comparable model complexity, achieves state-of-the-art performance on the synthetic and real-world image denoising tasks. We also provide code and models at https://github.com/gladzhang/Xformer.
format Preprint
id arxiv_https___arxiv_org_abs_2303_06440
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Xformer: Hybrid X-Shaped Transformer for Image Denoising
Zhang, Jiale
Zhang, Yulun
Gu, Jinjin
Dong, Jiahua
Kong, Linghe
Yang, Xiaokang
Computer Vision and Pattern Recognition
In this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two types of Transformer blocks. The spatial-wise Transformer block performs fine-grained local patches interactions across tokens defined by spatial dimension. The channel-wise Transformer block performs direct global context interactions across tokens defined by channel dimension. Based on the concurrent network structure, we design two branches to conduct these two interaction fashions. Within each branch, we employ an encoder-decoder architecture to capture multi-scale features. Besides, we propose the Bidirectional Connection Unit (BCU) to couple the learned representations from these two branches while providing enhanced information fusion. The joint designs make our Xformer powerful to conduct global information modeling in both spatial and channel dimensions. Extensive experiments show that Xformer, under the comparable model complexity, achieves state-of-the-art performance on the synthetic and real-world image denoising tasks. We also provide code and models at https://github.com/gladzhang/Xformer.
title Xformer: Hybrid X-Shaped Transformer for Image Denoising
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.06440