Neural Video Representation for Redundancy Reduction and Consistency Preservation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hayami, Taiga, Shindo, Takahiro, Akamatsu, Shunsuke, Watanabe, Hiroshi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909347349528576
author Hayami, Taiga
Shindo, Takahiro
Akamatsu, Shunsuke
Watanabe, Hiroshi
author_facet Hayami, Taiga
Shindo, Takahiro
Akamatsu, Shunsuke
Watanabe, Hiroshi
contents Implicit neural representation (INR) embed various signals into neural networks. They have gained attention in recent years because of their versatility in handling diverse signal types. In the context of video, INR achieves video compression by embedding video signals directly into networks and compressing them. Conventional methods either use an index that expresses the time of the frame or features extracted from individual frames as network inputs. The latter method provides greater expressive capability as the input is specific to each video. However, the features extracted from frames often contain redundancy, which contradicts the purpose of video compression. Additionally, such redundancies make it challenging to accurately reconstruct high-frequency components in the frames. To address these problems, we focus on separating the high-frequency and low-frequency components of the reconstructed frame. We propose a video representation method that generates both the high-frequency and low-frequency components of the frame, using features extracted from the high-frequency components and temporal information, respectively. Experimental results demonstrate that our method outperforms the existing HNeRV method, achieving superior results in 96 percent of the videos.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18497
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neural Video Representation for Redundancy Reduction and Consistency Preservation
Hayami, Taiga
Shindo, Takahiro
Akamatsu, Shunsuke
Watanabe, Hiroshi
Computer Vision and Pattern Recognition
Implicit neural representation (INR) embed various signals into neural networks. They have gained attention in recent years because of their versatility in handling diverse signal types. In the context of video, INR achieves video compression by embedding video signals directly into networks and compressing them. Conventional methods either use an index that expresses the time of the frame or features extracted from individual frames as network inputs. The latter method provides greater expressive capability as the input is specific to each video. However, the features extracted from frames often contain redundancy, which contradicts the purpose of video compression. Additionally, such redundancies make it challenging to accurately reconstruct high-frequency components in the frames. To address these problems, we focus on separating the high-frequency and low-frequency components of the reconstructed frame. We propose a video representation method that generates both the high-frequency and low-frequency components of the frame, using features extracted from the high-frequency components and temporal information, respectively. Experimental results demonstrate that our method outperforms the existing HNeRV method, achieving superior results in 96 percent of the videos.
title Neural Video Representation for Redundancy Reduction and Consistency Preservation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.18497