MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Phung, Huu-Tai, Gao, Zong-Lin, Yao, Yi-Chen, Ho, Kuan-Wei, Chen, Yi-Hsin, Lin, Yu-Hsiang, Gnutti, Alessandro, Peng, Wen-Hsiao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915553956855808
author Phung, Huu-Tai
Gao, Zong-Lin
Yao, Yi-Chen
Ho, Kuan-Wei
Chen, Yi-Hsin
Lin, Yu-Hsiang
Gnutti, Alessandro
Peng, Wen-Hsiao
author_facet Phung, Huu-Tai
Gao, Zong-Lin
Yao, Yi-Chen
Ho, Kuan-Wei
Chen, Yi-Hsin
Lin, Yu-Hsiang
Gnutti, Alessandro
Peng, Wen-Hsiao
contents This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and short-term reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional video coding offer superior coding performance. However, the need to store and access a large amount of implicit contextual information extracted from past decoded frames in decoding a video frame poses a challenge due to excessive memory access. Our MH-LVC overcomes this issue by storing multiple long- and short-term reference frames but limiting the number of reference frames used at a time for temporal prediction to two. Our decoded frame buffer management allows the encoder to flexibly utilize the long-term key frames to mitigate temporal cascading errors and the short-term reference frames to minimize prediction errors. Moreover, our buffering scheme enables the temporal prediction structure to be adapted to individual input videos. While this flexibility is common in traditional video codecs, it has not been fully explored for learned video codecs. Extensive experiments show that the proposed method outperforms VTM-17.0 under the low-delay B configuration in terms of PSNR-RGB across commonly used test datasets, and performs comparably to the state-of-the-art learned codecs (e.g.~DCVC-FM) while requiring less decoded frame buffer and similar decoding time.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding
Phung, Huu-Tai
Gao, Zong-Lin
Yao, Yi-Chen
Ho, Kuan-Wei
Chen, Yi-Hsin
Lin, Yu-Hsiang
Gnutti, Alessandro
Peng, Wen-Hsiao
Image and Video Processing
This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and short-term reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional video coding offer superior coding performance. However, the need to store and access a large amount of implicit contextual information extracted from past decoded frames in decoding a video frame poses a challenge due to excessive memory access. Our MH-LVC overcomes this issue by storing multiple long- and short-term reference frames but limiting the number of reference frames used at a time for temporal prediction to two. Our decoded frame buffer management allows the encoder to flexibly utilize the long-term key frames to mitigate temporal cascading errors and the short-term reference frames to minimize prediction errors. Moreover, our buffering scheme enables the temporal prediction structure to be adapted to individual input videos. While this flexibility is common in traditional video codecs, it has not been fully explored for learned video codecs. Extensive experiments show that the proposed method outperforms VTM-17.0 under the low-delay B configuration in terms of PSNR-RGB across commonly used test datasets, and performs comparably to the state-of-the-art learned codecs (e.g.~DCVC-FM) while requiring less decoded frame buffer and similar decoding time.
title MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding
topic Image and Video Processing
url https://arxiv.org/abs/2510.12479