MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jun, Zhang, Xinfeng, Tang, Lv, Jiang, JunHao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915350048669696
author Zhu, Jun
Zhang, Xinfeng
Tang, Lv
Jiang, JunHao
author_facet Zhu, Jun
Zhang, Xinfeng
Tang, Lv
Jiang, JunHao
contents Implicit Neural representations (INRs) have emerged as a promising approach for video compression, and have achieved comparable performance to the state-of-the-art codecs such as H.266/VVC. However, existing INR-based methods struggle to effectively represent detail-intensive and fast-changing video content. This limitation mainly stems from the underutilization of internal network features and the absence of video-specific considerations in network design. To address these challenges, we propose a multi-scale feature fusion framework, MSNeRV, for neural video representation. In the encoding stage, we enhance temporal consistency by employing temporal windows, and divide the video into multiple Groups of Pictures (GoPs), where a GoP-level grid is used for background representation. Additionally, we design a multi-scale spatial decoder with a scale-adaptive loss function to integrate multi-resolution and multi-frequency information. To further improve feature extraction, we introduce a multi-scale feature block that fully leverages hidden features. We evaluate MSNeRV on HEVC ClassB and UVG datasets for video representation and compression. Experimental results demonstrate that our model exhibits superior representation capability among INR-based approaches and surpasses VTM-23.7 (Random Access) in dynamic scenarios in terms of compression efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15276
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
Zhu, Jun
Zhang, Xinfeng
Tang, Lv
Jiang, JunHao
Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
Implicit Neural representations (INRs) have emerged as a promising approach for video compression, and have achieved comparable performance to the state-of-the-art codecs such as H.266/VVC. However, existing INR-based methods struggle to effectively represent detail-intensive and fast-changing video content. This limitation mainly stems from the underutilization of internal network features and the absence of video-specific considerations in network design. To address these challenges, we propose a multi-scale feature fusion framework, MSNeRV, for neural video representation. In the encoding stage, we enhance temporal consistency by employing temporal windows, and divide the video into multiple Groups of Pictures (GoPs), where a GoP-level grid is used for background representation. Additionally, we design a multi-scale spatial decoder with a scale-adaptive loss function to integrate multi-resolution and multi-frequency information. To further improve feature extraction, we introduce a multi-scale feature block that fully leverages hidden features. We evaluate MSNeRV on HEVC ClassB and UVG datasets for video representation and compression. Experimental results demonstrate that our model exhibits superior representation capability among INR-based approaches and surpasses VTM-23.7 (Random Access) in dynamic scenarios in terms of compression efficiency.
title MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
topic Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2506.15276