BVI-CR: A Multi-View Human Dataset for Volumetric Video Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Ge, Azzarelli, Adrian, Kwan, Ho Man, Anantrasirichai, Nantheera, Zhang, Fan, Moolan-Feroze, Oliver, Bull, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911046167429120
author Gao, Ge
Azzarelli, Adrian
Kwan, Ho Man
Anantrasirichai, Nantheera
Zhang, Fan
Moolan-Feroze, Oliver
Bull, David
author_facet Gao, Ge
Azzarelli, Adrian
Kwan, Ho Man
Anantrasirichai, Nantheera
Zhang, Fan
Moolan-Feroze, Oliver
Bull, David
contents The advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38\% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at \url{https://github.com/fan-aaron-zhang/bvi-cr}.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11199
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BVI-CR: A Multi-View Human Dataset for Volumetric Video Compression
Gao, Ge
Azzarelli, Adrian
Kwan, Ho Man
Anantrasirichai, Nantheera
Zhang, Fan
Moolan-Feroze, Oliver
Bull, David
Computer Vision and Pattern Recognition
Image and Video Processing
The advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38\% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at \url{https://github.com/fan-aaron-zhang/bvi-cr}.
title BVI-CR: A Multi-View Human Dataset for Volumetric Video Compression
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2411.11199