Direct Motion Models for Assessing Generated Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Allen, Kelsey, Doersch, Carl, Zhou, Guangyao, Suhail, Mohammed, Driess, Danny, Rocco, Ignacio, Rubanova, Yulia, Kipf, Thomas, Sajjadi, Mehdi S. M., Murphy, Kevin, Carreira, Joao, van Steenkiste, Sjoerd
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915268141252608
author Allen, Kelsey
Doersch, Carl
Zhou, Guangyao
Suhail, Mohammed
Driess, Danny
Rocco, Ignacio
Rubanova, Yulia
Kipf, Thomas
Sajjadi, Mehdi S. M.
Murphy, Kevin
Carreira, Joao
van Steenkiste, Sjoerd
author_facet Allen, Kelsey
Doersch, Carl
Zhou, Guangyao
Suhail, Mohammed
Driess, Danny
Rocco, Ignacio
Rubanova, Yulia
Kipf, Thomas
Sajjadi, Mehdi S. M.
Murphy, Kevin
Carreira, Joao
van Steenkiste, Sjoerd
contents A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond FVD by developing a metric which better measures plausible object interactions and motion. Our novel approach is based on auto-encoding point tracks and yields motion features that can be used to not only compare distributions of videos (as few as one generated and one ground truth, or as many as two datasets), but also for evaluating motion of single videos. We show that using point tracks instead of pixel reconstruction or action recognition features results in a metric which is markedly more sensitive to temporal distortions in synthetic data, and can predict human evaluations of temporal consistency and realism in generated videos obtained from open-source models better than a wide range of alternatives. We also show that by using a point track representation, we can spatiotemporally localize generative video inconsistencies, providing extra interpretability of generated video errors relative to prior work. An overview of the results and link to the code can be found on the project page: http://trajan-paper.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00209
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Direct Motion Models for Assessing Generated Videos
Allen, Kelsey
Doersch, Carl
Zhou, Guangyao
Suhail, Mohammed
Driess, Danny
Rocco, Ignacio
Rubanova, Yulia
Kipf, Thomas
Sajjadi, Mehdi S. M.
Murphy, Kevin
Carreira, Joao
van Steenkiste, Sjoerd
Computer Vision and Pattern Recognition
Machine Learning
A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond FVD by developing a metric which better measures plausible object interactions and motion. Our novel approach is based on auto-encoding point tracks and yields motion features that can be used to not only compare distributions of videos (as few as one generated and one ground truth, or as many as two datasets), but also for evaluating motion of single videos. We show that using point tracks instead of pixel reconstruction or action recognition features results in a metric which is markedly more sensitive to temporal distortions in synthetic data, and can predict human evaluations of temporal consistency and realism in generated videos obtained from open-source models better than a wide range of alternatives. We also show that by using a point track representation, we can spatiotemporally localize generative video inconsistencies, providing extra interpretability of generated video errors relative to prior work. An overview of the results and link to the code can be found on the project page: http://trajan-paper.github.io.
title Direct Motion Models for Assessing Generated Videos
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.00209