LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Murai, Shimon, Sun, Heming, Katto, Jiro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929598351015936
author Murai, Shimon
Sun, Heming
Katto, Jiro
author_facet Murai, Shimon
Sun, Heming
Katto, Jiro
contents Supported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achieve high compression rates and superior perceptual quality by using image captions as sub-information. This paper demonstrates that using a large multi-modal model (LMM), it is possible to generate captions and compress them within a single model. We also propose a novel semantic-perceptual-oriented fine-tuning method applicable to any LIC network, resulting in a 41.58\% improvement in LPIPS BD-rate compared to existing methods. Our implementation and pre-trained weights are available at https://github.com/tokkiwa/ImageTextCoding.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13033
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression
Murai, Shimon
Sun, Heming
Katto, Jiro
Image and Video Processing
Computer Vision and Pattern Recognition
Supported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achieve high compression rates and superior perceptual quality by using image captions as sub-information. This paper demonstrates that using a large multi-modal model (LMM), it is possible to generate captions and compress them within a single model. We also propose a novel semantic-perceptual-oriented fine-tuning method applicable to any LIC network, resulting in a 41.58\% improvement in LPIPS BD-rate compared to existing methods. Our implementation and pre-trained weights are available at https://github.com/tokkiwa/ImageTextCoding.
title LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13033