One-Step Diffusion-Based Image Compression with Semantic Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xue, Naifu, Jia, Zhaoyang, Li, Jiahao, Li, Bin, Zhang, Yuan, Lu, Yan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909926144606208
author Xue, Naifu
Jia, Zhaoyang
Li, Jiahao
Li, Bin
Zhang, Yuan
Lu, Yan
author_facet Xue, Naifu
Jia, Zhaoyang
Li, Jiahao
Li, Bin
Zhang, Yuan
Lu, Yan
contents While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasing latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based on this insight, we propose OneDC, a One-step Diffusion-based generative image Codec -- that integrates a latent compression module with a one-step diffusion generator. Recognizing the critical role of semantic guidance in one-step diffusion, we propose using the hyperprior as a semantic signal, overcoming the limitations of text prompts in representing complex visual content. To further enhance the semantic capability of the hyperprior, we introduce a semantic distillation mechanism that transfers knowledge from a pretrained generative tokenizer to the hyperprior codec. Additionally, we adopt a hybrid pixel- and latent-domain optimization to jointly enhance both reconstruction fidelity and perceptual realism. Extensive experiments demonstrate that OneDC achieves SOTA perceptual quality even with one-step generation, offering over 39% bitrate reduction and 20x faster decoding compared to prior multi-step diffusion-based codecs. Project: https://onedc-codec.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2505_16687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One-Step Diffusion-Based Image Compression with Semantic Distillation
Xue, Naifu
Jia, Zhaoyang
Li, Jiahao
Li, Bin
Zhang, Yuan
Lu, Yan
Computer Vision and Pattern Recognition
Image and Video Processing
While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasing latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based on this insight, we propose OneDC, a One-step Diffusion-based generative image Codec -- that integrates a latent compression module with a one-step diffusion generator. Recognizing the critical role of semantic guidance in one-step diffusion, we propose using the hyperprior as a semantic signal, overcoming the limitations of text prompts in representing complex visual content. To further enhance the semantic capability of the hyperprior, we introduce a semantic distillation mechanism that transfers knowledge from a pretrained generative tokenizer to the hyperprior codec. Additionally, we adopt a hybrid pixel- and latent-domain optimization to jointly enhance both reconstruction fidelity and perceptual realism. Extensive experiments demonstrate that OneDC achieves SOTA perceptual quality even with one-step generation, offering over 39% bitrate reduction and 20x faster decoding compared to prior multi-step diffusion-based codecs. Project: https://onedc-codec.github.io/
title One-Step Diffusion-Based Image Compression with Semantic Distillation
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2505.16687