Invertible Residual Rescaling Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jinmin, Dai, Tao, Zha, Yaohua, Luo, Yilu, Lu, Longfei, Chen, Bin, Wang, Zhi, Xia, Shu-Tao, Zhang, Jingyun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914793015738368
author Li, Jinmin
Dai, Tao
Zha, Yaohua
Luo, Yilu
Lu, Longfei
Chen, Bin
Wang, Zhi
Xia, Shu-Tao
Zhang, Jingyun
author_facet Li, Jinmin
Dai, Tao
Zha, Yaohua
Luo, Yilu
Lu, Longfei
Chen, Bin
Wang, Zhi
Xia, Shu-Tao
Zhang, Jingyun
contents Invertible Rescaling Networks (IRNs) and their variants have witnessed remarkable achievements in various image processing tasks like image rescaling. However, we observe that IRNs with deeper networks are difficult to train, thus hindering the representational ability of IRNs. To address this issue, we propose Invertible Residual Rescaling Models (IRRM) for image rescaling by learning a bijection between a high-resolution image and its low-resolution counterpart with a specific distribution. Specifically, we propose IRRM to build a deep network, which contains several Residual Downscaling Modules (RDMs) with long skip connections. Each RDM consists of several Invertible Residual Blocks (IRBs) with short connections. In this way, RDM allows rich low-frequency information to be bypassed by skip connections and forces models to focus on extracting high-frequency information from the image. Extensive experiments show that our IRRM performs significantly better than other state-of-the-art methods with much fewer parameters and complexity. Particularly, our IRRM has respectively PSNR gains of at least 0.3 dB over HCFlow and IRN in the x4 rescaling while only using 60% parameters and 50% FLOPs. The code will be available at https://github.com/THU-Kingmin/IRRM.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02945
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Invertible Residual Rescaling Models
Li, Jinmin
Dai, Tao
Zha, Yaohua
Luo, Yilu
Lu, Longfei
Chen, Bin
Wang, Zhi
Xia, Shu-Tao
Zhang, Jingyun
Computer Vision and Pattern Recognition
Invertible Rescaling Networks (IRNs) and their variants have witnessed remarkable achievements in various image processing tasks like image rescaling. However, we observe that IRNs with deeper networks are difficult to train, thus hindering the representational ability of IRNs. To address this issue, we propose Invertible Residual Rescaling Models (IRRM) for image rescaling by learning a bijection between a high-resolution image and its low-resolution counterpart with a specific distribution. Specifically, we propose IRRM to build a deep network, which contains several Residual Downscaling Modules (RDMs) with long skip connections. Each RDM consists of several Invertible Residual Blocks (IRBs) with short connections. In this way, RDM allows rich low-frequency information to be bypassed by skip connections and forces models to focus on extracting high-frequency information from the image. Extensive experiments show that our IRRM performs significantly better than other state-of-the-art methods with much fewer parameters and complexity. Particularly, our IRRM has respectively PSNR gains of at least 0.3 dB over HCFlow and IRN in the x4 rescaling while only using 60% parameters and 50% FLOPs. The code will be available at https://github.com/THU-Kingmin/IRRM.
title Invertible Residual Rescaling Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.02945