Restoration-Oriented Video Frame Interpolation with Region-Distinguishable Priors from SAM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Yan, Xu, Xiaogang, Lin, Yingqi, Wu, Jiafei, Liu, Zhe, Yang, Ming-Hsuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914535766491136
author Han, Yan
Xu, Xiaogang
Lin, Yingqi
Wu, Jiafei
Liu, Zhe
Yang, Ming-Hsuan
author_facet Han, Yan
Xu, Xiaogang
Lin, Yingqi
Wu, Jiafei
Liu, Zhe
Yang, Ming-Hsuan
contents In existing restoration-oriented Video Frame Interpolation (VFI) approaches, the motion estimation between neighboring frames plays a crucial role. However, the estimation accuracy in existing methods remains a challenge, primarily due to the inherent ambiguity in identifying corresponding areas in adjacent frames for interpolation. Therefore, enhancing accuracy by distinguishing different regions before motion estimation is of utmost importance. In this paper, we introduce a novel solution involving the utilization of open-world segmentation models, e.g., SAM2 (Segment Anything Model2) for frames, to derive Region-Distinguishable Priors (RDPs) in different frames. These RDPs are represented as spatial-varying Gaussian mixtures, distinguishing an arbitrary number of areas with a unified modality. RDPs can be integrated into existing motion-based VFI methods to enhance features for motion estimation, facilitated by our designed play-and-plug Hierarchical Region-aware Feature Fusion Module (HRFFM). HRFFM incorporates RDP into various hierarchical stages of VFI's encoder, using RDP-guided Feature Normalization (RDPFN) in a residual learning manner. With HRFFM and RDP, the features within VFI's encoder exhibit similar representations for matched regions in neighboring frames, thus improving the synthesis of intermediate frames. Extensive experiments demonstrate that HRFFM consistently enhances VFI performance across various scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2312_15868
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Restoration-Oriented Video Frame Interpolation with Region-Distinguishable Priors from SAM
Han, Yan
Xu, Xiaogang
Lin, Yingqi
Wu, Jiafei
Liu, Zhe
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
In existing restoration-oriented Video Frame Interpolation (VFI) approaches, the motion estimation between neighboring frames plays a crucial role. However, the estimation accuracy in existing methods remains a challenge, primarily due to the inherent ambiguity in identifying corresponding areas in adjacent frames for interpolation. Therefore, enhancing accuracy by distinguishing different regions before motion estimation is of utmost importance. In this paper, we introduce a novel solution involving the utilization of open-world segmentation models, e.g., SAM2 (Segment Anything Model2) for frames, to derive Region-Distinguishable Priors (RDPs) in different frames. These RDPs are represented as spatial-varying Gaussian mixtures, distinguishing an arbitrary number of areas with a unified modality. RDPs can be integrated into existing motion-based VFI methods to enhance features for motion estimation, facilitated by our designed play-and-plug Hierarchical Region-aware Feature Fusion Module (HRFFM). HRFFM incorporates RDP into various hierarchical stages of VFI's encoder, using RDP-guided Feature Normalization (RDPFN) in a residual learning manner. With HRFFM and RDP, the features within VFI's encoder exhibit similar representations for matched regions in neighboring frames, thus improving the synthesis of intermediate frames. Extensive experiments demonstrate that HRFFM consistently enhances VFI performance across various scenes.
title Restoration-Oriented Video Frame Interpolation with Region-Distinguishable Priors from SAM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.15868