On Disentangled Training for Nonlinear Transform in Learned Image Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Han, Li, Shaohui, Dai, Wenrui, Cao, Maida, Kan, Nuowen, Li, Chenglin, Zou, Junni, Xiong, Hongkai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913692136767488
author Li, Han
Li, Shaohui
Dai, Wenrui
Cao, Maida
Kan, Nuowen
Li, Chenglin
Zou, Junni
Xiong, Hongkai
author_facet Li, Han
Li, Shaohui
Dai, Wenrui
Cao, Maida
Kan, Nuowen
Li, Chenglin
Zou, Junni
Xiong, Hongkai
contents Learned image compression (LIC) has demonstrated superior rate-distortion (R-D) performance compared to traditional codecs, but is challenged by training inefficiency that could incur more than two weeks to train a state-of-the-art model from scratch. Existing LIC methods overlook the slow convergence caused by compacting energy in learning nonlinear transforms. In this paper, we first reveal that such energy compaction consists of two components, i.e., feature decorrelation and uneven energy modulation. On such basis, we propose a linear auxiliary transform (AuxT) to disentangle energy compaction in training nonlinear transforms. The proposed AuxT obtains coarse approximation to achieve efficient energy compaction such that distribution fitting with the nonlinear transforms can be simplified to fine details. We then develop wavelet-based linear shortcuts (WLSs) for AuxT that leverages wavelet-based downsampling and orthogonal linear projection for feature decorrelation and subband-aware scaling for
format Preprint
id arxiv_https___arxiv_org_abs_2501_13751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Disentangled Training for Nonlinear Transform in Learned Image Compression
Li, Han
Li, Shaohui
Dai, Wenrui
Cao, Maida
Kan, Nuowen
Li, Chenglin
Zou, Junni
Xiong, Hongkai
Image and Video Processing
Computer Vision and Pattern Recognition
Learned image compression (LIC) has demonstrated superior rate-distortion (R-D) performance compared to traditional codecs, but is challenged by training inefficiency that could incur more than two weeks to train a state-of-the-art model from scratch. Existing LIC methods overlook the slow convergence caused by compacting energy in learning nonlinear transforms. In this paper, we first reveal that such energy compaction consists of two components, i.e., feature decorrelation and uneven energy modulation. On such basis, we propose a linear auxiliary transform (AuxT) to disentangle energy compaction in training nonlinear transforms. The proposed AuxT obtains coarse approximation to achieve efficient energy compaction such that distribution fitting with the nonlinear transforms can be simplified to fine details. We then develop wavelet-based linear shortcuts (WLSs) for AuxT that leverages wavelet-based downsampling and orthogonal linear projection for feature decorrelation and subband-aware scaling for
title On Disentangled Training for Nonlinear Transform in Learned Image Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.13751