VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vicente, Noel José Rodrigues, Lehner, Enrique, Villar-Corrales, Angel, Nogga, Jan, Behnke, Sven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913910320267264
author Vicente, Noel José Rodrigues
Lehner, Enrique
Villar-Corrales, Angel
Nogga, Jan
Behnke, Sven
author_facet Vicente, Noel José Rodrigues
Lehner, Enrique
Villar-Corrales, Angel
Nogga, Jan
Behnke, Sven
contents Understanding and predicting video content is essential for planning and reasoning in dynamic environments. Despite advancements, unsupervised learning of object representations and dynamics remains challenging. We present VideoPCDNet, an unsupervised framework for object-centric video decomposition and prediction. Our model uses frequency-domain phase correlation techniques to recursively parse videos into object components, which are represented as transformed versions of learned object prototypes, enabling accurate and interpretable tracking. By explicitly modeling object motion through a combination of frequency domain operations and lightweight learned modules, VideoPCDNet enables accurate unsupervised object tracking and prediction of future video frames. In our experiments, we demonstrate that VideoPCDNet outperforms multiple object-centric baseline models for unsupervised tracking and prediction on several synthetic datasets, while learning interpretable object and motion representations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
Vicente, Noel José Rodrigues
Lehner, Enrique
Villar-Corrales, Angel
Nogga, Jan
Behnke, Sven
Computer Vision and Pattern Recognition
Artificial Intelligence
Understanding and predicting video content is essential for planning and reasoning in dynamic environments. Despite advancements, unsupervised learning of object representations and dynamics remains challenging. We present VideoPCDNet, an unsupervised framework for object-centric video decomposition and prediction. Our model uses frequency-domain phase correlation techniques to recursively parse videos into object components, which are represented as transformed versions of learned object prototypes, enabling accurate and interpretable tracking. By explicitly modeling object motion through a combination of frequency domain operations and lightweight learned modules, VideoPCDNet enables accurate unsupervised object tracking and prediction of future video frames. In our experiments, we demonstrate that VideoPCDNet outperforms multiple object-centric baseline models for unsupervised tracking and prediction on several synthetic datasets, while learning interpretable object and motion representations.
title VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.19621