High-Fidelity Image Compression with Score-based Generative Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoogeboom, Emiel, Agustsson, Eirikur, Mentzer, Fabian, Versari, Luca, Toderici, George, Theis, Lucas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910358017409024
author Hoogeboom, Emiel
Agustsson, Eirikur
Mentzer, Fabian
Versari, Luca
Toderici, George
Theis, Lucas
author_facet Hoogeboom, Emiel
Agustsson, Eirikur
Mentzer, Fabian
Versari, Luca
Toderici, George
Theis, Lucas
contents Despite the tremendous success of diffusion generative models in text-to-image generation, replicating this success in the domain of image compression has proven difficult. In this paper, we demonstrate that diffusion can significantly improve perceptual quality at a given bit-rate, outperforming state-of-the-art approaches PO-ELIC and HiFiC as measured by FID score. This is achieved using a simple but theoretically motivated two-stage approach combining an autoencoder targeting MSE followed by a further score-based decoder. However, as we will show, implementation details matter and the optimal design decisions can differ greatly from typical text-to-image models.
format Preprint
id arxiv_https___arxiv_org_abs_2305_18231
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle High-Fidelity Image Compression with Score-based Generative Models
Hoogeboom, Emiel
Agustsson, Eirikur
Mentzer, Fabian
Versari, Luca
Toderici, George
Theis, Lucas
Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
Despite the tremendous success of diffusion generative models in text-to-image generation, replicating this success in the domain of image compression has proven difficult. In this paper, we demonstrate that diffusion can significantly improve perceptual quality at a given bit-rate, outperforming state-of-the-art approaches PO-ELIC and HiFiC as measured by FID score. This is achieved using a simple but theoretically motivated two-stage approach combining an autoencoder targeting MSE followed by a further score-based decoder. However, as we will show, implementation details matter and the optimal design decisions can differ greatly from typical text-to-image models.
title High-Fidelity Image Compression with Score-based Generative Models
topic Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2305.18231