Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914981532925952 |
|---|---|
| author | Meo, Cristian Nakano, Akihiro Lică, Mircea Didolkar, Aniket Suzuki, Masahiro Goyal, Anirudh Zhang, Mengmi Dauwels, Justin Matsuo, Yutaka Bengio, Yoshua |
| author_facet | Meo, Cristian Nakano, Akihiro Lică, Mircea Didolkar, Aniket Suzuki, Masahiro Goyal, Anirudh Zhang, Mengmi Dauwels, Justin Matsuo, Yutaka Bengio, Yoshua |
| contents | Unsupervised object-centric learning from videos is a promising approach towards learning compositional representations that can be applied to various downstream tasks, such as prediction and reasoning. Recently, it was shown that pretrained Vision Transformers (ViTs) can be useful to learn object-centric representations on real-world video datasets. However, while these approaches succeed at extracting objects from the scenes, the slot-based representations fail to maintain temporal consistency across consecutive frames in a video, i.e. the mapping of objects to slots changes across the video. To address this, we introduce Conditional Autoregressive Slot Attention (CA-SA), a framework that enhances the temporal consistency of extracted object-centric representations in video-centric vision tasks. Leveraging an autoregressive prior network to condition representations on previous timesteps and a novel consistency loss function, CA-SA predicts future slot representations and imposes consistency across frames. We present qualitative and quantitative results showing that our proposed method outperforms the considered baselines on downstream tasks, such as video prediction and visual question-answering tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_15728 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases Meo, Cristian Nakano, Akihiro Lică, Mircea Didolkar, Aniket Suzuki, Masahiro Goyal, Anirudh Zhang, Mengmi Dauwels, Justin Matsuo, Yutaka Bengio, Yoshua Computer Vision and Pattern Recognition Machine Learning Unsupervised object-centric learning from videos is a promising approach towards learning compositional representations that can be applied to various downstream tasks, such as prediction and reasoning. Recently, it was shown that pretrained Vision Transformers (ViTs) can be useful to learn object-centric representations on real-world video datasets. However, while these approaches succeed at extracting objects from the scenes, the slot-based representations fail to maintain temporal consistency across consecutive frames in a video, i.e. the mapping of objects to slots changes across the video. To address this, we introduce Conditional Autoregressive Slot Attention (CA-SA), a framework that enhances the temporal consistency of extracted object-centric representations in video-centric vision tasks. Leveraging an autoregressive prior network to condition representations on previous timesteps and a novel consistency loss function, CA-SA predicts future slot representations and imposes consistency across frames. We present qualitative and quantitative results showing that our proposed method outperforms the considered baselines on downstream tasks, such as video prediction and visual question-answering tasks. |
| title | Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2410.15728 |