CoD: A Diffusion Foundation Model for Image Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Zhaoyang, Zheng, Zihan, Xue, Naifu, Li, Jiahao, Li, Bin, Guo, Zongyu, Zhang, Xiaoyi, Li, Houqiang, Lu, Yan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911519067865088
author Jia, Zhaoyang
Zheng, Zihan
Xue, Naifu
Li, Jiahao
Li, Bin
Guo, Zongyu
Zhang, Xiaoyi
Li, Houqiang
Lu, Yan
author_facet Jia, Zhaoyang
Zheng, Zihan
Xue, Naifu
Li, Jiahao
Li, Bin
Guo, Zongyu
Zhang, Xiaoyi
Li, Houqiang
Lu, Yan
contents Existing diffusion codecs typically build on text-to-image diffusion foundation models like Stable Diffusion. However, text conditioning is suboptimal from a compression perspective, hindering the potential of downstream diffusion codecs, particularly at ultra-low bitrates. To address it, we introduce \textbf{CoD}, the first \textbf{Co}mpression-oriented \textbf{D}iffusion foundation model, trained from scratch to enable end-to-end optimization of both compression and generation. CoD is not a fixed codec but a general foundation model designed for various diffusion-based codecs. It offers several advantages: \textbf{High compression efficiency}, replacing Stable Diffusion with CoD in downstream codecs like DiffC achieves SOTA results, especially at ultra-low bitrates (e.g., 0.0039 bpp); \textbf{Low-cost and reproducible training}, 300$\times$ faster training than Stable Diffusion ($\sim$ 20 vs. $\sim$ 6,250 A100 GPU days) on entirely open image-only datasets; \textbf{Providing new insights}, e.g., We find pixel-space diffusion can achieve VTM-level PSNR with high perceptual quality and can outperform GAN-based codecs using fewer parameters. We hope CoD lays the foundation for future diffusion codec research. Codes are released at https://github.com/microsoft/GenCodec/tree/main/CoD.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoD: A Diffusion Foundation Model for Image Compression
Jia, Zhaoyang
Zheng, Zihan
Xue, Naifu
Li, Jiahao
Li, Bin
Guo, Zongyu
Zhang, Xiaoyi
Li, Houqiang
Lu, Yan
Computer Vision and Pattern Recognition
Existing diffusion codecs typically build on text-to-image diffusion foundation models like Stable Diffusion. However, text conditioning is suboptimal from a compression perspective, hindering the potential of downstream diffusion codecs, particularly at ultra-low bitrates. To address it, we introduce \textbf{CoD}, the first \textbf{Co}mpression-oriented \textbf{D}iffusion foundation model, trained from scratch to enable end-to-end optimization of both compression and generation. CoD is not a fixed codec but a general foundation model designed for various diffusion-based codecs. It offers several advantages: \textbf{High compression efficiency}, replacing Stable Diffusion with CoD in downstream codecs like DiffC achieves SOTA results, especially at ultra-low bitrates (e.g., 0.0039 bpp); \textbf{Low-cost and reproducible training}, 300$\times$ faster training than Stable Diffusion ($\sim$ 20 vs. $\sim$ 6,250 A100 GPU days) on entirely open image-only datasets; \textbf{Providing new insights}, e.g., We find pixel-space diffusion can achieve VTM-level PSNR with high perceptual quality and can outperform GAN-based codecs using fewer parameters. We hope CoD lays the foundation for future diffusion codec research. Codes are released at https://github.com/microsoft/GenCodec/tree/main/CoD.
title CoD: A Diffusion Foundation Model for Image Compression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18706