A Multi-Grid Implicit Neural Representation for Multi-View Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ling, Qingyue, Cheng, Zhengxue, Feng, Donghui, Wang, Shen, Zhu, Chen, Lu, Guo, Sun, Heming, Katto, Jiro, Song, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916959022481408
author Ling, Qingyue
Cheng, Zhengxue
Feng, Donghui
Wang, Shen
Zhu, Chen
Lu, Guo
Sun, Heming
Katto, Jiro
Song, Li
author_facet Ling, Qingyue
Cheng, Zhengxue
Feng, Donghui
Wang, Shen
Zhu, Chen
Lu, Guo
Sun, Heming
Katto, Jiro
Song, Li
contents Multi-view videos are becoming widely used in different fields, but their high resolution and multi-camera shooting raise significant challenges for storage and transmission. In this paper, we propose MV-MGINR, a multi-grid implicit neural representation for multi-view videos. It combines a time-indexed grid, a view-indexed grid and an integrated time and view grid. The first two grids capture common representative contents across each view and time axis respectively, and the latter one captures local details under specific view and time. Then, a synthesis net is used to upsample the multi-grid latents and generate reconstructed frames. Additionally, a motion-aware loss is introduced to enhance the reconstruction quality of moving regions. The proposed framework effectively integrates the common and local features of multi-view videos, ultimately achieving high-quality reconstruction. Compared with MPEG immersive video test model TMIV, MV-MGINR achieves bitrate savings of 72.3% while maintaining the same PSNR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multi-Grid Implicit Neural Representation for Multi-View Videos
Ling, Qingyue
Cheng, Zhengxue
Feng, Donghui
Wang, Shen
Zhu, Chen
Lu, Guo
Sun, Heming
Katto, Jiro
Song, Li
Image and Video Processing
Multi-view videos are becoming widely used in different fields, but their high resolution and multi-camera shooting raise significant challenges for storage and transmission. In this paper, we propose MV-MGINR, a multi-grid implicit neural representation for multi-view videos. It combines a time-indexed grid, a view-indexed grid and an integrated time and view grid. The first two grids capture common representative contents across each view and time axis respectively, and the latter one captures local details under specific view and time. Then, a synthesis net is used to upsample the multi-grid latents and generate reconstructed frames. Additionally, a motion-aware loss is introduced to enhance the reconstruction quality of moving regions. The proposed framework effectively integrates the common and local features of multi-view videos, ultimately achieving high-quality reconstruction. Compared with MPEG immersive video test model TMIV, MV-MGINR achieves bitrate savings of 72.3% while maintaining the same PSNR.
title A Multi-Grid Implicit Neural Representation for Multi-View Videos
topic Image and Video Processing
url https://arxiv.org/abs/2509.16706