Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914306786852864 |
|---|---|
| author | Zhang, Maojun Wu, Haotian Jin, Richeng Gunduz, Deniz Mikolajczyk, Krystian |
| author_facet | Zhang, Maojun Wu, Haotian Jin, Richeng Gunduz, Deniz Mikolajczyk, Krystian |
| contents | Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm for video compression by leveraging high-level semantic understanding and powerful visual synthesis. This paper propose a video compression framework that integrates generative priors to drastically reduce bit-rate while maintaining reconstruction fidelity. Specifically, our method compresses high-level semantic representations of the video, then uses a conditional diffusion model to reconstruct frames from these semantics. To further improve compression, we characterize motion information with global camera trajectories and foreground segmentation: background motion is compactly represented by camera pose parameters while foreground dynamics by sparse segmentation masks. This allows for significantly boosts compression efficiency, enabling descent video reconstruction at extremely low bit-rates. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_05201 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance Zhang, Maojun Wu, Haotian Jin, Richeng Gunduz, Deniz Mikolajczyk, Krystian Image and Video Processing Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm for video compression by leveraging high-level semantic understanding and powerful visual synthesis. This paper propose a video compression framework that integrates generative priors to drastically reduce bit-rate while maintaining reconstruction fidelity. Specifically, our method compresses high-level semantic representations of the video, then uses a conditional diffusion model to reconstruct frames from these semantics. To further improve compression, we characterize motion information with global camera trajectories and foreground segmentation: background motion is compactly represented by camera pose parameters while foreground dynamics by sparse segmentation masks. This allows for significantly boosts compression efficiency, enabling descent video reconstruction at extremely low bit-rates. |
| title | Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance |
| topic | Image and Video Processing |
| url | https://arxiv.org/abs/2602.05201 |