OSCAR: One-Step Diffusion Codec Across Multiple Bit-rates

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guo, Jinpei, Ji, Yifei, Chen, Zheng, Liu, Kai, Liu, Min, Rao, Wang, Li, Wenbo, Guo, Yong, Zhang, Yulun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918163497615360
author Guo, Jinpei
Ji, Yifei
Chen, Zheng
Liu, Kai
Liu, Min
Rao, Wang
Li, Wenbo
Guo, Yong
Zhang, Yulun
author_facet Guo, Jinpei
Ji, Yifei
Chen, Zheng
Liu, Kai
Liu, Min
Rao, Wang
Li, Wenbo
Guo, Yong
Zhang, Yulun
contents Pretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising from random noise, guided by compressed latent representations. While these approaches have achieved high reconstruction quality, their multi-step sampling process incurs substantial computational overhead. Moreover, they typically require training separate models for different compression bit-rates, leading to significant training and storage costs. To address these challenges, we propose a one-step diffusion codec across multiple bit-rates. termed OSCAR. Specifically, our method views compressed latents as noisy variants of the original latents, where the level of distortion depends on the bit-rate. This perspective allows them to be modeled as intermediate states along a diffusion trajectory. By establishing a mapping from the compression bit-rate to a pseudo diffusion timestep, we condition a single generative model to support reconstructions at multiple bit-rates. Meanwhile, we argue that the compressed latents retain rich structural information, thereby making one-step denoising feasible. Thus, OSCAR replaces iterative sampling with a single denoising pass, significantly improving inference efficiency. Extensive experiments demonstrate that OSCAR achieves superior performance in both quantitative and visual quality metrics. The code and models are available at https://github.com/jp-guo/OSCAR.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OSCAR: One-Step Diffusion Codec Across Multiple Bit-rates
Guo, Jinpei
Ji, Yifei
Chen, Zheng
Liu, Kai
Liu, Min
Rao, Wang
Li, Wenbo
Guo, Yong
Zhang, Yulun
Image and Video Processing
Computer Vision and Pattern Recognition
Pretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising from random noise, guided by compressed latent representations. While these approaches have achieved high reconstruction quality, their multi-step sampling process incurs substantial computational overhead. Moreover, they typically require training separate models for different compression bit-rates, leading to significant training and storage costs. To address these challenges, we propose a one-step diffusion codec across multiple bit-rates. termed OSCAR. Specifically, our method views compressed latents as noisy variants of the original latents, where the level of distortion depends on the bit-rate. This perspective allows them to be modeled as intermediate states along a diffusion trajectory. By establishing a mapping from the compression bit-rate to a pseudo diffusion timestep, we condition a single generative model to support reconstructions at multiple bit-rates. Meanwhile, we argue that the compressed latents retain rich structural information, thereby making one-step denoising feasible. Thus, OSCAR replaces iterative sampling with a single denoising pass, significantly improving inference efficiency. Extensive experiments demonstrate that OSCAR achieves superior performance in both quantitative and visual quality metrics. The code and models are available at https://github.com/jp-guo/OSCAR.
title OSCAR: One-Step Diffusion Codec Across Multiple Bit-rates
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16091