Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918494435540992 |
|---|---|
| author | Li, Zhiyuan Zhao, Rongzhen Yang, Wenyan Zhao, Wenshuai Marttinen, Pekka Pajarinen, Joni |
| author_facet | Li, Zhiyuan Zhao, Rongzhen Yang, Wenyan Zhao, Wenshuai Marttinen, Pekka Pajarinen, Joni |
| contents | The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive approximations of discrete correspondence problems. Modern self-supervised vision backbones already encode instance-discriminative features that distinguish objects reliably. Exploiting these features eliminates the need for learned temporal prediction. We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. Slots initialize from salient regions in frozen backbone features. Frame-to-frame identity is maintained through Hungarian matching on slot representations. The approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on MOVi-D, MOVi-E, and YouTube-VIS. Project page: https://magenta-sherbet-85b101.netlify.app/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_03650 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence Li, Zhiyuan Zhao, Rongzhen Yang, Wenyan Zhao, Wenshuai Marttinen, Pekka Pajarinen, Joni Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive approximations of discrete correspondence problems. Modern self-supervised vision backbones already encode instance-discriminative features that distinguish objects reliably. Exploiting these features eliminates the need for learned temporal prediction. We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. Slots initialize from salient regions in frozen backbone features. Frame-to-frame identity is maintained through Hungarian matching on slot representations. The approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on MOVi-D, MOVi-E, and YouTube-VIS. Project page: https://magenta-sherbet-85b101.netlify.app/ |
| title | Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2605.03650 |