Fast Training-free Perceptual Image Compression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Ziran, Xu, Tongda, Huang, Minye, He, Dailan, Ge, Xingtong, Zhang, Xinjie, Li, Ling, Wang, Yan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915350795255808
author Zhu, Ziran
Xu, Tongda
Huang, Minye
He, Dailan
Ge, Xingtong
Zhang, Xinjie
Li, Ling
Wang, Yan
author_facet Zhu, Ziran
Xu, Tongda
Huang, Minye
He, Dailan
Ge, Xingtong
Zhang, Xinjie
Li, Ling
Wang, Yan
contents Training-free perceptual image codec adopt pre-trained unconditional generative model during decoding to avoid training new conditional generative model. However, they heavily rely on diffusion inversion or sample communication, which take 1 min to intractable amount of time to decode a single image. In this paper, we propose a training-free algorithm that improves the perceptual quality of any existing codec with theoretical guarantee. We further propose different implementations for optimal perceptual quality when decoding time budget is $\approx 0.1$s, $0.1-10$s and $\ge 10$s. Our approach: 1). improves the decoding time of training-free codec from 1 min to $0.1-10$s with comparable perceptual quality. 2). can be applied to non-differentiable codec such as VTM. 3). can be used to improve previous perceptual codecs, such as MS-ILLM. 4). can easily achieve perception-distortion trade-off. Empirically, we show that our approach successfully improves the perceptual quality of ELIC, VTM and MS-ILLM with fast decoding. Our approach achieves comparable FID to previous training-free codec with significantly less decoding time. And our approach still outperforms previous conditional generative model based codecs such as HiFiC and MS-ILLM in terms of FID. The source code is provided in the supplementary material.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fast Training-free Perceptual Image Compression
Zhu, Ziran
Xu, Tongda
Huang, Minye
He, Dailan
Ge, Xingtong
Zhang, Xinjie
Li, Ling
Wang, Yan
Image and Video Processing
Computer Vision and Pattern Recognition
Training-free perceptual image codec adopt pre-trained unconditional generative model during decoding to avoid training new conditional generative model. However, they heavily rely on diffusion inversion or sample communication, which take 1 min to intractable amount of time to decode a single image. In this paper, we propose a training-free algorithm that improves the perceptual quality of any existing codec with theoretical guarantee. We further propose different implementations for optimal perceptual quality when decoding time budget is $\approx 0.1$s, $0.1-10$s and $\ge 10$s. Our approach: 1). improves the decoding time of training-free codec from 1 min to $0.1-10$s with comparable perceptual quality. 2). can be applied to non-differentiable codec such as VTM. 3). can be used to improve previous perceptual codecs, such as MS-ILLM. 4). can easily achieve perception-distortion trade-off. Empirically, we show that our approach successfully improves the perceptual quality of ELIC, VTM and MS-ILLM with fast decoding. Our approach achieves comparable FID to previous training-free codec with significantly less decoding time. And our approach still outperforms previous conditional generative model based codecs such as HiFiC and MS-ILLM in terms of FID. The source code is provided in the supplementary material.
title Fast Training-free Perceptual Image Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.16102