Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiyuan, Zhao, Rongzhen, Yang, Wenyan, Zhao, Wenshuai, Marttinen, Pekka, Pajarinen, Joni
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918494435540992
author Li, Zhiyuan
Zhao, Rongzhen
Yang, Wenyan
Zhao, Wenshuai
Marttinen, Pekka
Pajarinen, Joni
author_facet Li, Zhiyuan
Zhao, Rongzhen
Yang, Wenyan
Zhao, Wenshuai
Marttinen, Pekka
Pajarinen, Joni
contents The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive approximations of discrete correspondence problems. Modern self-supervised vision backbones already encode instance-discriminative features that distinguish objects reliably. Exploiting these features eliminates the need for learned temporal prediction. We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. Slots initialize from salient regions in frozen backbone features. Frame-to-frame identity is maintained through Hungarian matching on slot representations. The approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on MOVi-D, MOVi-E, and YouTube-VIS. Project page: https://magenta-sherbet-85b101.netlify.app/
format Preprint
id arxiv_https___arxiv_org_abs_2605_03650
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Li, Zhiyuan
Zhao, Rongzhen
Yang, Wenyan
Zhao, Wenshuai
Marttinen, Pekka
Pajarinen, Joni
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive approximations of discrete correspondence problems. Modern self-supervised vision backbones already encode instance-discriminative features that distinguish objects reliably. Exploiting these features eliminates the need for learned temporal prediction. We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. Slots initialize from salient regions in frozen backbone features. Frame-to-frame identity is maintained through Hungarian matching on slot representations. The approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on MOVi-D, MOVi-E, and YouTube-VIS. Project page: https://magenta-sherbet-85b101.netlify.app/
title Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.03650