UniMIC: Towards Universal Multi-modality Perceptual Image Compression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Yixin, Li, Xin, Pan, Xiaohan, Feng, Runsen, Guo, Zongyu, Lu, Yiting, Ren, Yulin, Chen, Zhibo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913602135392256
author Gao, Yixin
Li, Xin
Pan, Xiaohan
Feng, Runsen
Guo, Zongyu
Lu, Yiting
Ren, Yulin
Chen, Zhibo
author_facet Gao, Yixin
Li, Xin
Pan, Xiaohan
Feng, Runsen
Guo, Zongyu
Lu, Yiting
Ren, Yulin
Chen, Zhibo
contents We present UniMIC, a universal multi-modality image compression framework, intending to unify the rate-distortion-perception (RDP) optimization for multiple image codecs simultaneously through excavating cross-modality generative priors. Unlike most existing works that need to design and optimize image codecs from scratch, our UniMIC introduces the visual codec repository, which incorporates amounts of representative image codecs and directly uses them as the basic codecs for various practical applications. Moreover, we propose multi-grained textual coding, where variable-length content prompt and compression prompt are designed and encoded to assist the perceptual reconstruction through the multi-modality conditional generation. In particular, a universal perception compensator is proposed to improve the perception quality of decoded images from all basic codecs at the decoder side by reusing text-assisted diffusion priors from stable diffusion. With the cooperation of the above three strategies, our UniMIC achieves a significant improvement of RDP optimization for different compression codecs, e.g., traditional and learnable codecs, and different compression costs, e.g., ultra-low bitrates. The code will be available in https://github.com/Amygyx/UniMIC .
format Preprint
id arxiv_https___arxiv_org_abs_2412_04912
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniMIC: Towards Universal Multi-modality Perceptual Image Compression
Gao, Yixin
Li, Xin
Pan, Xiaohan
Feng, Runsen
Guo, Zongyu
Lu, Yiting
Ren, Yulin
Chen, Zhibo
Image and Video Processing
Computer Vision and Pattern Recognition
We present UniMIC, a universal multi-modality image compression framework, intending to unify the rate-distortion-perception (RDP) optimization for multiple image codecs simultaneously through excavating cross-modality generative priors. Unlike most existing works that need to design and optimize image codecs from scratch, our UniMIC introduces the visual codec repository, which incorporates amounts of representative image codecs and directly uses them as the basic codecs for various practical applications. Moreover, we propose multi-grained textual coding, where variable-length content prompt and compression prompt are designed and encoded to assist the perceptual reconstruction through the multi-modality conditional generation. In particular, a universal perception compensator is proposed to improve the perception quality of decoded images from all basic codecs at the decoder side by reusing text-assisted diffusion priors from stable diffusion. With the cooperation of the above three strategies, our UniMIC achieves a significant improvement of RDP optimization for different compression codecs, e.g., traditional and learnable codecs, and different compression costs, e.g., ultra-low bitrates. The code will be available in https://github.com/Amygyx/UniMIC .
title UniMIC: Towards Universal Multi-modality Perceptual Image Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04912