MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guo, Jialong, liu, Ke, Yao, Jiangchao, Wang, Zhihua, Bu, Jiajun, Wang, Haishuai
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909458512216064
author Guo, Jialong
liu, Ke
Yao, Jiangchao
Wang, Zhihua
Bu, Jiajun
Wang, Haishuai
author_facet Guo, Jialong
liu, Ke
Yao, Jiangchao
Wang, Zhihua
Bu, Jiajun
Wang, Haishuai
contents Neural Representations for Videos (NeRV) has emerged as a promising implicit neural representation (INR) approach for video analysis, which represents videos as neural networks with frame indexes as inputs. However, NeRV-based methods are time-consuming when adapting to a large number of diverse videos, as each video requires a separate NeRV model to be trained from scratch. In addition, NeRV-based methods spatially require generating a high-dimension signal (i.e., an entire image) from the input of a low-dimension timestamp, and a video typically consists of tens of frames temporally that have a minor change between adjacent frames. To improve the efficiency of video representation, we propose Meta Neural Representations for Videos, named MetaNeRV, a novel framework for fast NeRV representation for unseen videos. MetaNeRV leverages a meta-learning framework to learn an optimal parameter initialization, which serves as a good starting point for adapting to new videos. To address the unique spatial and temporal characteristics of video modality, we further introduce spatial-temporal guidance to improve the representation capabilities of MetaNeRV. Specifically, the spatial guidance with a multi-resolution loss aims to capture the information from different resolution stages, and the temporal guidance with an effective progressive learning strategy could gradually refine the number of fitted frames during the meta-learning process. Extensive experiments conducted on multiple datasets demonstrate the superiority of MetaNeRV for video representations and video compression.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance
Guo, Jialong
liu, Ke
Yao, Jiangchao
Wang, Zhihua
Bu, Jiajun
Wang, Haishuai
Computer Vision and Pattern Recognition
Neural Representations for Videos (NeRV) has emerged as a promising implicit neural representation (INR) approach for video analysis, which represents videos as neural networks with frame indexes as inputs. However, NeRV-based methods are time-consuming when adapting to a large number of diverse videos, as each video requires a separate NeRV model to be trained from scratch. In addition, NeRV-based methods spatially require generating a high-dimension signal (i.e., an entire image) from the input of a low-dimension timestamp, and a video typically consists of tens of frames temporally that have a minor change between adjacent frames. To improve the efficiency of video representation, we propose Meta Neural Representations for Videos, named MetaNeRV, a novel framework for fast NeRV representation for unseen videos. MetaNeRV leverages a meta-learning framework to learn an optimal parameter initialization, which serves as a good starting point for adapting to new videos. To address the unique spatial and temporal characteristics of video modality, we further introduce spatial-temporal guidance to improve the representation capabilities of MetaNeRV. Specifically, the spatial guidance with a multi-resolution loss aims to capture the information from different resolution stages, and the temporal guidance with an effective progressive learning strategy could gradually refine the number of fitted frames during the meta-learning process. Extensive experiments conducted on multiple datasets demonstrate the superiority of MetaNeRV for video representations and video compression.
title MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.02427