Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911883627331584 |
|---|---|
| author | Lee, Hagyeong Kim, Minkyu Kim, Jun-Hyuk Kim, Seungeon Oh, Dokwan Lee, Jaeho |
| author_facet | Lee, Hagyeong Kim, Minkyu Kim, Jun-Hyuk Kim, Seungeon Oh, Dokwan Lee, Jaeho |
| contents | Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image compression algorithm that achieves both high perceptual and pixel-wise fidelity. In particular, we propose a compression framework that leverages text information mainly by text-adaptive encoding and training with joint image-text loss. By doing so, we avoid decoding based on text-guided generative models -- known for high generative diversity -- and effectively utilize the semantic information of text at a global level. Experimental results on various datasets show that our method can achieve high pixel-level and perceptual quality, with either human- or machine-generated captions. In particular, our method outperforms all baselines in terms of LPIPS, with some room for even more improvements when we use more carefully generated captions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_02944 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity Lee, Hagyeong Kim, Minkyu Kim, Jun-Hyuk Kim, Seungeon Oh, Dokwan Lee, Jaeho Computer Vision and Pattern Recognition Machine Learning Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image compression algorithm that achieves both high perceptual and pixel-wise fidelity. In particular, we propose a compression framework that leverages text information mainly by text-adaptive encoding and training with joint image-text loss. By doing so, we avoid decoding based on text-guided generative models -- known for high generative diversity -- and effectively utilize the semantic information of text at a global level. Experimental results on various datasets show that our method can achieve high pixel-level and perceptual quality, with either human- or machine-generated captions. In particular, our method outperforms all baselines in terms of LPIPS, with some room for even more improvements when we use more carefully generated captions. |
| title | Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2403.02944 |