Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zadaianchuk, Andrii, Seitzer, Maximilian, Martius, Georg
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929276811476992
author Zadaianchuk, Andrii
Seitzer, Maximilian
Martius, Georg
author_facet Zadaianchuk, Andrii
Seitzer, Maximilian
Martius, Georg
contents Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains. Recently, it was shown that the reconstruction of pre-trained self-supervised features leads to object-centric representations on unconstrained real-world image datasets. Building on this approach, we propose a novel way to use such pre-trained features in the form of a temporal feature similarity loss. This loss encodes semantic and temporal correlations between image patches and is a natural way to introduce a motion bias for object discovery. We demonstrate that this loss leads to state-of-the-art performance on the challenging synthetic MOVi datasets. When used in combination with the feature reconstruction loss, our model is the first object-centric video model that scales to unconstrained video datasets such as YouTube-VIS.
format Preprint
id arxiv_https___arxiv_org_abs_2306_04829
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
Zadaianchuk, Andrii
Seitzer, Maximilian
Martius, Georg
Computer Vision and Pattern Recognition
Machine Learning
Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains. Recently, it was shown that the reconstruction of pre-trained self-supervised features leads to object-centric representations on unconstrained real-world image datasets. Building on this approach, we propose a novel way to use such pre-trained features in the form of a temporal feature similarity loss. This loss encodes semantic and temporal correlations between image patches and is a natural way to introduce a motion bias for object discovery. We demonstrate that this loss leads to state-of-the-art performance on the challenging synthetic MOVi datasets. When used in combination with the feature reconstruction loss, our model is the first object-centric video model that scales to unconstrained video datasets such as YouTube-VIS.
title Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2306.04829