Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shrivastava, Gaurav, Shrivastava, Abhinav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910734252769280
author Shrivastava, Gaurav
Shrivastava, Abhinav
author_facet Shrivastava, Gaurav
Shrivastava, Abhinav
contents Diffusion models have made significant strides in image generation, mastering tasks such as unconditional image synthesis, text-image translation, and image-to-image conversions. However, their capability falls short in the realm of video prediction, mainly because they treat videos as a collection of independent images, relying on external constraints such as temporal attention mechanisms to enforce temporal coherence. In our paper, we introduce a novel model class, that treats video as a continuous multi-dimensional process rather than a series of discrete frames. We also report a reduction of 75\% sampling steps required to sample a new frame thus making our framework more efficient during the inference time. Through extensive experimentation, we establish state-of-the-art performance in video prediction, validated on benchmark datasets including KTH, BAIR, Human3.6M, and UCF101. Navigate to the project page https://www.cs.umd.edu/~gauravsh/cvp/supp/website.html for video results.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04929
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
Shrivastava, Gaurav
Shrivastava, Abhinav
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Diffusion models have made significant strides in image generation, mastering tasks such as unconditional image synthesis, text-image translation, and image-to-image conversions. However, their capability falls short in the realm of video prediction, mainly because they treat videos as a collection of independent images, relying on external constraints such as temporal attention mechanisms to enforce temporal coherence. In our paper, we introduce a novel model class, that treats video as a continuous multi-dimensional process rather than a series of discrete frames. We also report a reduction of 75\% sampling steps required to sample a new frame thus making our framework more efficient during the inference time. Through extensive experimentation, we establish state-of-the-art performance in video prediction, validated on benchmark datasets including KTH, BAIR, Human3.6M, and UCF101. Navigate to the project page https://www.cs.umd.edu/~gauravsh/cvp/supp/website.html for video results.
title Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.04929