Unified Arbitrary-Time Video Frame Interpolation and Prediction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jin, Xin, Wu, Longhai, Chen, Jie, Cho, Ilhyun, Hahm, Cheul-Hee
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916642498281472
author Jin, Xin
Wu, Longhai
Chen, Jie
Cho, Ilhyun
Hahm, Cheul-Hee
author_facet Jin, Xin
Wu, Longhai
Chen, Jie
Cho, Ilhyun
Hahm, Cheul-Hee
contents Video frame interpolation and prediction aim to synthesize frames in-between and subsequent to existing frames, respectively. Despite being closely-related, these two tasks are traditionally studied with different model architectures, or same architecture but individually trained weights. Furthermore, while arbitrary-time interpolation has been extensively studied, the value of arbitrary-time prediction has been largely overlooked. In this work, we present uniVIP - unified arbitrary-time Video Interpolation and Prediction. Technically, we firstly extend an interpolation-only network for arbitrary-time interpolation and prediction, with a special input channel for task (interpolation or prediction) encoding. Then, we show how to train a unified model on common triplet frames. Our uniVIP provides competitive results for video interpolation, and outperforms existing state-of-the-arts for video prediction. Codes will be available at: https://github.com/srcn-ivl/uniVIP
format Preprint
id arxiv_https___arxiv_org_abs_2503_02316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified Arbitrary-Time Video Frame Interpolation and Prediction
Jin, Xin
Wu, Longhai
Chen, Jie
Cho, Ilhyun
Hahm, Cheul-Hee
Computer Vision and Pattern Recognition
Video frame interpolation and prediction aim to synthesize frames in-between and subsequent to existing frames, respectively. Despite being closely-related, these two tasks are traditionally studied with different model architectures, or same architecture but individually trained weights. Furthermore, while arbitrary-time interpolation has been extensively studied, the value of arbitrary-time prediction has been largely overlooked. In this work, we present uniVIP - unified arbitrary-time Video Interpolation and Prediction. Technically, we firstly extend an interpolation-only network for arbitrary-time interpolation and prediction, with a special input channel for task (interpolation or prediction) encoding. Then, we show how to train a unified model on common triplet frames. Our uniVIP provides competitive results for video interpolation, and outperforms existing state-of-the-arts for video prediction. Codes will be available at: https://github.com/srcn-ivl/uniVIP
title Unified Arbitrary-Time Video Frame Interpolation and Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.02316