On Disentangled Training for Nonlinear Transform in Learned Image Compression

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Han, Li, Shaohui, Dai, Wenrui, Cao, Maida, Kan, Nuowen, Li, Chenglin, Zou, Junni, Xiong, Hongkai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913692136767488
author Li, Han
Li, Shaohui
Dai, Wenrui
Cao, Maida
Kan, Nuowen
Li, Chenglin
Zou, Junni
Xiong, Hongkai
author_facet Li, Han
Li, Shaohui
Dai, Wenrui
Cao, Maida
Kan, Nuowen
Li, Chenglin
Zou, Junni
Xiong, Hongkai
contents Learned image compression (LIC) has demonstrated superior rate-distortion (R-D) performance compared to traditional codecs, but is challenged by training inefficiency that could incur more than two weeks to train a state-of-the-art model from scratch. Existing LIC methods overlook the slow convergence caused by compacting energy in learning nonlinear transforms. In this paper, we first reveal that such energy compaction consists of two components, i.e., feature decorrelation and uneven energy modulation. On such basis, we propose a linear auxiliary transform (AuxT) to disentangle energy compaction in training nonlinear transforms. The proposed AuxT obtains coarse approximation to achieve efficient energy compaction such that distribution fitting with the nonlinear transforms can be simplified to fine details. We then develop wavelet-based linear shortcuts (WLSs) for AuxT that leverages wavelet-based downsampling and orthogonal linear projection for feature decorrelation and subband-aware scaling for
format Preprint
id arxiv_https___arxiv_org_abs_2501_13751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Disentangled Training for Nonlinear Transform in Learned Image Compression
Li, Han
Li, Shaohui
Dai, Wenrui
Cao, Maida
Kan, Nuowen
Li, Chenglin
Zou, Junni
Xiong, Hongkai
Image and Video Processing
Computer Vision and Pattern Recognition
Learned image compression (LIC) has demonstrated superior rate-distortion (R-D) performance compared to traditional codecs, but is challenged by training inefficiency that could incur more than two weeks to train a state-of-the-art model from scratch. Existing LIC methods overlook the slow convergence caused by compacting energy in learning nonlinear transforms. In this paper, we first reveal that such energy compaction consists of two components, i.e., feature decorrelation and uneven energy modulation. On such basis, we propose a linear auxiliary transform (AuxT) to disentangle energy compaction in training nonlinear transforms. The proposed AuxT obtains coarse approximation to achieve efficient energy compaction such that distribution fitting with the nonlinear transforms can be simplified to fine details. We then develop wavelet-based linear shortcuts (WLSs) for AuxT that leverages wavelet-based downsampling and orthogonal linear projection for feature decorrelation and subband-aware scaling for
title On Disentangled Training for Nonlinear Transform in Learned Image Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.13751