Video prediction using score-based conditional density estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fiquet, Pierre-Étienne H., Simoncelli, Eero P.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913570295382016
author Fiquet, Pierre-Étienne H.
Simoncelli, Eero P.
author_facet Fiquet, Pierre-Étienne H.
Simoncelli, Eero P.
contents Temporal prediction is inherently uncertain, but representing the ambiguity in natural image sequences is a challenging high-dimensional probabilistic inference problem. For natural scenes, the curse of dimensionality renders explicit density estimation statistically and computationally intractable. Here, we describe an implicit regression-based framework for learning and sampling the conditional density of the next frame in a video given previous observed frames. We show that sequence-to-image deep networks trained on a simple resilience-to-noise objective function extract adaptive representations for temporal prediction. Synthetic experiments demonstrate that this score-based framework can handle occlusion boundaries: unlike classical methods that average over bifurcating temporal trajectories, it chooses among likely trajectories, selecting more probable options with higher frequency. Furthermore, analysis of networks trained on natural image sequences reveals that the representation automatically weights predictive evidence by its reliability, which is a hallmark of statistical inference
format Preprint
id arxiv_https___arxiv_org_abs_2411_00842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video prediction using score-based conditional density estimation
Fiquet, Pierre-Étienne H.
Simoncelli, Eero P.
Computer Vision and Pattern Recognition
Machine Learning
Temporal prediction is inherently uncertain, but representing the ambiguity in natural image sequences is a challenging high-dimensional probabilistic inference problem. For natural scenes, the curse of dimensionality renders explicit density estimation statistically and computationally intractable. Here, we describe an implicit regression-based framework for learning and sampling the conditional density of the next frame in a video given previous observed frames. We show that sequence-to-image deep networks trained on a simple resilience-to-noise objective function extract adaptive representations for temporal prediction. Synthetic experiments demonstrate that this score-based framework can handle occlusion boundaries: unlike classical methods that average over bifurcating temporal trajectories, it chooses among likely trajectories, selecting more probable options with higher frequency. Furthermore, analysis of networks trained on natural image sequences reveals that the representation automatically weights predictive evidence by its reliability, which is a hallmark of statistical inference
title Video prediction using score-based conditional density estimation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.00842