On the Rate-Distortion-Complexity Trade-offs of Neural Video Coding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Yi-Hsin, Ho, Kuan-Wei, Benjak, Martin, Ostermann, Jörn, Peng, Wen-Hsiao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910634542628864
author Chen, Yi-Hsin
Ho, Kuan-Wei
Benjak, Martin
Ostermann, Jörn
Peng, Wen-Hsiao
author_facet Chen, Yi-Hsin
Ho, Kuan-Wei
Benjak, Martin
Ostermann, Jörn
Peng, Wen-Hsiao
contents This paper aims to delve into the rate-distortion-complexity trade-offs of modern neural video coding. Recent years have witnessed much research effort being focused on exploring the full potential of neural video coding. Conditional autoencoders have emerged as the mainstream approach to efficient neural video coding. The central theme of conditional autoencoders is to leverage both spatial and temporal information for better conditional coding. However, a recent study indicates that conditional coding may suffer from information bottlenecks, potentially performing worse than traditional residual coding. To address this issue, recent conditional coding methods incorporate a large number of high-resolution features as the condition signal, leading to a considerable increase in the number of multiply-accumulate operations, memory footprint, and model size. Taking DCVC as the common code base, we investigate how the newly proposed conditional residual coding, an emerging new school of thought, and its variants may strike a better balance among rate, distortion, and complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03898
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Rate-Distortion-Complexity Trade-offs of Neural Video Coding
Chen, Yi-Hsin
Ho, Kuan-Wei
Benjak, Martin
Ostermann, Jörn
Peng, Wen-Hsiao
Image and Video Processing
This paper aims to delve into the rate-distortion-complexity trade-offs of modern neural video coding. Recent years have witnessed much research effort being focused on exploring the full potential of neural video coding. Conditional autoencoders have emerged as the mainstream approach to efficient neural video coding. The central theme of conditional autoencoders is to leverage both spatial and temporal information for better conditional coding. However, a recent study indicates that conditional coding may suffer from information bottlenecks, potentially performing worse than traditional residual coding. To address this issue, recent conditional coding methods incorporate a large number of high-resolution features as the condition signal, leading to a considerable increase in the number of multiply-accumulate operations, memory footprint, and model size. Taking DCVC as the common code base, we investigate how the newly proposed conditional residual coding, an emerging new school of thought, and its variants may strike a better balance among rate, distortion, and complexity.
title On the Rate-Distortion-Complexity Trade-offs of Neural Video Coding
topic Image and Video Processing
url https://arxiv.org/abs/2410.03898