Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meo, Cristian, Nakano, Akihiro, Lică, Mircea, Didolkar, Aniket, Suzuki, Masahiro, Goyal, Anirudh, Zhang, Mengmi, Dauwels, Justin, Matsuo, Yutaka, Bengio, Yoshua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914981532925952
author Meo, Cristian
Nakano, Akihiro
Lică, Mircea
Didolkar, Aniket
Suzuki, Masahiro
Goyal, Anirudh
Zhang, Mengmi
Dauwels, Justin
Matsuo, Yutaka
Bengio, Yoshua
author_facet Meo, Cristian
Nakano, Akihiro
Lică, Mircea
Didolkar, Aniket
Suzuki, Masahiro
Goyal, Anirudh
Zhang, Mengmi
Dauwels, Justin
Matsuo, Yutaka
Bengio, Yoshua
contents Unsupervised object-centric learning from videos is a promising approach towards learning compositional representations that can be applied to various downstream tasks, such as prediction and reasoning. Recently, it was shown that pretrained Vision Transformers (ViTs) can be useful to learn object-centric representations on real-world video datasets. However, while these approaches succeed at extracting objects from the scenes, the slot-based representations fail to maintain temporal consistency across consecutive frames in a video, i.e. the mapping of objects to slots changes across the video. To address this, we introduce Conditional Autoregressive Slot Attention (CA-SA), a framework that enhances the temporal consistency of extracted object-centric representations in video-centric vision tasks. Leveraging an autoregressive prior network to condition representations on previous timesteps and a novel consistency loss function, CA-SA predicts future slot representations and imposes consistency across frames. We present qualitative and quantitative results showing that our proposed method outperforms the considered baselines on downstream tasks, such as video prediction and visual question-answering tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15728
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
Meo, Cristian
Nakano, Akihiro
Lică, Mircea
Didolkar, Aniket
Suzuki, Masahiro
Goyal, Anirudh
Zhang, Mengmi
Dauwels, Justin
Matsuo, Yutaka
Bengio, Yoshua
Computer Vision and Pattern Recognition
Machine Learning
Unsupervised object-centric learning from videos is a promising approach towards learning compositional representations that can be applied to various downstream tasks, such as prediction and reasoning. Recently, it was shown that pretrained Vision Transformers (ViTs) can be useful to learn object-centric representations on real-world video datasets. However, while these approaches succeed at extracting objects from the scenes, the slot-based representations fail to maintain temporal consistency across consecutive frames in a video, i.e. the mapping of objects to slots changes across the video. To address this, we introduce Conditional Autoregressive Slot Attention (CA-SA), a framework that enhances the temporal consistency of extracted object-centric representations in video-centric vision tasks. Leveraging an autoregressive prior network to condition representations on previous timesteps and a novel consistency loss function, CA-SA predicts future slot representations and imposes consistency across frames. We present qualitative and quantitative results showing that our proposed method outperforms the considered baselines on downstream tasks, such as video prediction and visual question-answering tasks.
title Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.15728