Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Hagyeong, Kim, Minkyu, Kim, Jun-Hyuk, Kim, Seungeon, Oh, Dokwan, Lee, Jaeho
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911883627331584
author Lee, Hagyeong
Kim, Minkyu
Kim, Jun-Hyuk
Kim, Seungeon
Oh, Dokwan
Lee, Jaeho
author_facet Lee, Hagyeong
Kim, Minkyu
Kim, Jun-Hyuk
Kim, Seungeon
Oh, Dokwan
Lee, Jaeho
contents Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image compression algorithm that achieves both high perceptual and pixel-wise fidelity. In particular, we propose a compression framework that leverages text information mainly by text-adaptive encoding and training with joint image-text loss. By doing so, we avoid decoding based on text-guided generative models -- known for high generative diversity -- and effectively utilize the semantic information of text at a global level. Experimental results on various datasets show that our method can achieve high pixel-level and perceptual quality, with either human- or machine-generated captions. In particular, our method outperforms all baselines in terms of LPIPS, with some room for even more improvements when we use more carefully generated captions.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02944
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity
Lee, Hagyeong
Kim, Minkyu
Kim, Jun-Hyuk
Kim, Seungeon
Oh, Dokwan
Lee, Jaeho
Computer Vision and Pattern Recognition
Machine Learning
Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image compression algorithm that achieves both high perceptual and pixel-wise fidelity. In particular, we propose a compression framework that leverages text information mainly by text-adaptive encoding and training with joint image-text loss. By doing so, we avoid decoding based on text-guided generative models -- known for high generative diversity -- and effectively utilize the semantic information of text at a global level. Experimental results on various datasets show that our method can achieve high pixel-level and perceptual quality, with either human- or machine-generated captions. In particular, our method outperforms all baselines in terms of LPIPS, with some room for even more improvements when we use more carefully generated captions.
title Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.02944