Immersive Video Compression using Implicit Neural Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwan, Ho Man, Zhang, Fan, Gower, Andrew, Bull, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915028264812544
author Kwan, Ho Man
Zhang, Fan
Gower, Andrew
Bull, David
author_facet Kwan, Ho Man
Zhang, Fan
Gower, Andrew
Bull, David
contents Recent work on implicit neural representations (INRs) has evidenced their potential for efficiently representing and encoding conventional video content. In this paper we, for the first time, extend their application to immersive (multi-view) videos, by proposing MV-HiNeRV, a new INR-based immersive video codec. MV-HiNeRV is an enhanced version of a state-of-the-art INR-based video codec, HiNeRV, which was developed for single-view video compression. We have modified the model to learn a different group of feature grids for each view, and share the learnt network parameters among all views. This enables the model to effectively exploit the spatio-temporal and the inter-view redundancy that exists within multi-view videos. The proposed codec was used to compress multi-view texture and depth video sequences in the MPEG Immersive Video (MIV) Common Test Conditions, and tested against the MIV Test model (TMIV) that uses the VVenC video codec. The results demonstrate the superior performance of MV-HiNeRV, with significant coding gains (up to 72.33\%) over TMIV. The implementation of MV-HiNeRV is published for further development and evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Immersive Video Compression using Implicit Neural Representations
Kwan, Ho Man
Zhang, Fan
Gower, Andrew
Bull, David
Image and Video Processing
Computer Vision and Pattern Recognition
Recent work on implicit neural representations (INRs) has evidenced their potential for efficiently representing and encoding conventional video content. In this paper we, for the first time, extend their application to immersive (multi-view) videos, by proposing MV-HiNeRV, a new INR-based immersive video codec. MV-HiNeRV is an enhanced version of a state-of-the-art INR-based video codec, HiNeRV, which was developed for single-view video compression. We have modified the model to learn a different group of feature grids for each view, and share the learnt network parameters among all views. This enables the model to effectively exploit the spatio-temporal and the inter-view redundancy that exists within multi-view videos. The proposed codec was used to compress multi-view texture and depth video sequences in the MPEG Immersive Video (MIV) Common Test Conditions, and tested against the MIV Test model (TMIV) that uses the VVenC video codec. The results demonstrate the superior performance of MV-HiNeRV, with significant coding gains (up to 72.33\%) over TMIV. The implementation of MV-HiNeRV is published for further development and evaluation.
title Immersive Video Compression using Implicit Neural Representations
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.01596