Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Maojun, Wu, Haotian, Jin, Richeng, Gunduz, Deniz, Mikolajczyk, Krystian
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914306786852864
author Zhang, Maojun
Wu, Haotian
Jin, Richeng
Gunduz, Deniz
Mikolajczyk, Krystian
author_facet Zhang, Maojun
Wu, Haotian
Jin, Richeng
Gunduz, Deniz
Mikolajczyk, Krystian
contents Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm for video compression by leveraging high-level semantic understanding and powerful visual synthesis. This paper propose a video compression framework that integrates generative priors to drastically reduce bit-rate while maintaining reconstruction fidelity. Specifically, our method compresses high-level semantic representations of the video, then uses a conditional diffusion model to reconstruct frames from these semantics. To further improve compression, we characterize motion information with global camera trajectories and foreground segmentation: background motion is compactly represented by camera pose parameters while foreground dynamics by sparse segmentation masks. This allows for significantly boosts compression efficiency, enabling descent video reconstruction at extremely low bit-rates.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05201
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance
Zhang, Maojun
Wu, Haotian
Jin, Richeng
Gunduz, Deniz
Mikolajczyk, Krystian
Image and Video Processing
Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm for video compression by leveraging high-level semantic understanding and powerful visual synthesis. This paper propose a video compression framework that integrates generative priors to drastically reduce bit-rate while maintaining reconstruction fidelity. Specifically, our method compresses high-level semantic representations of the video, then uses a conditional diffusion model to reconstruct frames from these semantics. To further improve compression, we characterize motion information with global camera trajectories and foreground segmentation: background motion is compactly represented by camera pose parameters while foreground dynamics by sparse segmentation masks. This allows for significantly boosts compression efficiency, enabling descent video reconstruction at extremely low bit-rates.
title Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance
topic Image and Video Processing
url https://arxiv.org/abs/2602.05201