PICD: Versatile Perceptual Image Compression with Diffusion Rendering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Tongda, Li, Jiahao, Li, Bin, Wang, Yan, Zhang, Ya-Qin, Lu, Yan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909605926273024
author Xu, Tongda
Li, Jiahao
Li, Bin
Wang, Yan
Zhang, Ya-Qin
Lu, Yan
author_facet Xu, Tongda
Li, Jiahao
Li, Bin
Wang, Yan
Zhang, Ya-Qin
Lu, Yan
contents Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PICD: Versatile Perceptual Image Compression with Diffusion Rendering
Xu, Tongda
Li, Jiahao
Li, Bin
Wang, Yan
Zhang, Ya-Qin
Lu, Yan
Computer Vision and Pattern Recognition
Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images.
title PICD: Versatile Perceptual Image Compression with Diffusion Rendering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.05853