Multi-Scale Invertible Neural Network for Wide-Range Variable-Rate Learned Image Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Hanyue, Wu, Siqi, Li, Li, Zhou, Wengang, Li, Houqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916664455462912
author Tu, Hanyue
Wu, Siqi
Li, Li
Zhou, Wengang
Li, Houqiang
author_facet Tu, Hanyue
Wu, Siqi
Li, Li
Zhou, Wengang
Li, Houqiang
contents Autoencoder-based structures have dominated recent learned image compression methods. However, the inherent information loss associated with autoencoders limits their rate-distortion performance at high bit rates and restricts their flexibility of rate adaptation. In this paper, we present a variable-rate image compression model based on invertible transform to overcome these limitations. Specifically, we design a lightweight multi-scale invertible neural network, which bijectively maps the input image into multi-scale latent representations. To improve the compression efficiency, a multi-scale spatial-channel context model with extended gain units is devised to estimate the entropy of the latent representation from high to low levels. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared to existing variable-rate methods, and remains competitive with recent multi-model approaches. Notably, our method is the first learned image compression solution that outperforms VVC across a very wide range of bit rates using a single model, especially at high bit rates. The source code is available at https://github.com/hytu99/MSINN-VRLIC.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Scale Invertible Neural Network for Wide-Range Variable-Rate Learned Image Compression
Tu, Hanyue
Wu, Siqi
Li, Li
Zhou, Wengang
Li, Houqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Autoencoder-based structures have dominated recent learned image compression methods. However, the inherent information loss associated with autoencoders limits their rate-distortion performance at high bit rates and restricts their flexibility of rate adaptation. In this paper, we present a variable-rate image compression model based on invertible transform to overcome these limitations. Specifically, we design a lightweight multi-scale invertible neural network, which bijectively maps the input image into multi-scale latent representations. To improve the compression efficiency, a multi-scale spatial-channel context model with extended gain units is devised to estimate the entropy of the latent representation from high to low levels. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared to existing variable-rate methods, and remains competitive with recent multi-model approaches. Notably, our method is the first learned image compression solution that outperforms VVC across a very wide range of bit rates using a single model, especially at high bit rates. The source code is available at https://github.com/hytu99/MSINN-VRLIC.
title Multi-Scale Invertible Neural Network for Wide-Range Variable-Rate Learned Image Compression
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.21284