MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Chunyi, Lu, Guo, Feng, Donghui, Wu, Haoning, Zhang, Zicheng, Liu, Xiaohong, Zhai, Guangtao, Lin, Weisi, Zhang, Wenjun
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914758018465792
author Li, Chunyi
Lu, Guo
Feng, Donghui
Wu, Haoning
Zhang, Zicheng
Liu, Xiaohong
Zhai, Guangtao
Lin, Weisi
Zhang, Wenjun
author_facet Li, Chunyi
Lu, Guo
Feng, Donghui
Wu, Haoning
Zhang, Zicheng
Liu, Xiaohong
Zhai, Guangtao
Lin, Weisi
Zhang, Wenjun
contents With the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. In recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16749
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model
Li, Chunyi
Lu, Guo
Feng, Donghui
Wu, Haoning
Zhang, Zicheng
Liu, Xiaohong
Zhai, Guangtao
Lin, Weisi
Zhang, Wenjun
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
With the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. In recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC.
title MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
url https://arxiv.org/abs/2402.16749