Efficient Continuous Video Flow Model for Video Prediction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shrivastava, Gaurav, Shrivastava, Abhinav
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909419297570816
author Shrivastava, Gaurav
Shrivastava, Abhinav
author_facet Shrivastava, Gaurav
Shrivastava, Abhinav
contents Multi-step prediction models, such as diffusion and rectified flow models, have emerged as state-of-the-art solutions for generation tasks. However, these models exhibit higher latency in sampling new frames compared to single-step methods. This latency issue becomes a significant bottleneck when adapting such methods for video prediction tasks, given that a typical 60-second video comprises approximately 1.5K frames. In this paper, we propose a novel approach to modeling the multi-step process, aimed at alleviating latency constraints and facilitating the adaptation of such processes for video prediction tasks. Our approach not only reduces the number of sample steps required to predict the next frame but also minimizes computational demands by reducing the model size to one-third of the original size. We evaluate our method on standard video prediction datasets, including KTH, BAIR action robot, Human3.6M and UCF101, demonstrating its efficacy in achieving state-of-the-art performance on these benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05633
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Continuous Video Flow Model for Video Prediction
Shrivastava, Gaurav
Shrivastava, Abhinav
Computer Vision and Pattern Recognition
Multi-step prediction models, such as diffusion and rectified flow models, have emerged as state-of-the-art solutions for generation tasks. However, these models exhibit higher latency in sampling new frames compared to single-step methods. This latency issue becomes a significant bottleneck when adapting such methods for video prediction tasks, given that a typical 60-second video comprises approximately 1.5K frames. In this paper, we propose a novel approach to modeling the multi-step process, aimed at alleviating latency constraints and facilitating the adaptation of such processes for video prediction tasks. Our approach not only reduces the number of sample steps required to predict the next frame but also minimizes computational demands by reducing the model size to one-third of the original size. We evaluate our method on standard video prediction datasets, including KTH, BAIR action robot, Human3.6M and UCF101, demonstrating its efficacy in achieving state-of-the-art performance on these benchmarks.
title Efficient Continuous Video Flow Model for Video Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.05633