VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Qiang, Zhong, Houqiang, Zheng, Zihan, Zhang, Xiaoyun, Cheng, Zhengxue, Song, Li, Zhai, Guangtao, Wang, Yanfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915065417957376
author Hu, Qiang
Zhong, Houqiang
Zheng, Zihan
Zhang, Xiaoyun
Cheng, Zhengxue
Song, Li
Zhai, Guangtao
Wang, Yanfeng
author_facet Hu, Qiang
Zhong, Houqiang
Zheng, Zihan
Zhang, Xiaoyun
Cheng, Zhengxue
Song, Li
Zhai, Guangtao
Wang, Yanfeng
contents Neural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11362
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression
Hu, Qiang
Zhong, Houqiang
Zheng, Zihan
Zhang, Xiaoyun
Cheng, Zhengxue
Song, Li
Zhai, Guangtao
Wang, Yanfeng
Image and Video Processing
Computer Vision and Pattern Recognition
Neural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets.
title VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11362