Identifiable Object Representations under Spatial Ambiguities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kori, Avinash, Toni, Francesca, Glocker, Ben
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909643719049216
author Kori, Avinash
Toni, Francesca
Glocker, Ben
author_facet Kori, Avinash
Toni, Francesca
Glocker, Ben
contents Modular object-centric representations are essential for *human-like reasoning* but are challenging to obtain under spatial ambiguities, *e.g. due to occlusions and view ambiguities*. However, addressing challenges presents both theoretical and practical difficulties. We introduce a novel multi-view probabilistic approach that aggregates view-specific slots to capture *invariant content* information while simultaneously learning disentangled global *viewpoint-level* information. Unlike prior single-view methods, our approach resolves spatial ambiguities, provides theoretical guarantees for identifiability, and requires *no viewpoint annotations*. Extensive experiments on standard benchmarks and novel complex datasets validate our method's robustness and scalability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Identifiable Object Representations under Spatial Ambiguities
Kori, Avinash
Toni, Francesca
Glocker, Ben
Machine Learning
Computer Vision and Pattern Recognition
Modular object-centric representations are essential for *human-like reasoning* but are challenging to obtain under spatial ambiguities, *e.g. due to occlusions and view ambiguities*. However, addressing challenges presents both theoretical and practical difficulties. We introduce a novel multi-view probabilistic approach that aggregates view-specific slots to capture *invariant content* information while simultaneously learning disentangled global *viewpoint-level* information. Unlike prior single-view methods, our approach resolves spatial ambiguities, provides theoretical guarantees for identifiability, and requires *no viewpoint annotations*. Extensive experiments on standard benchmarks and novel complex datasets validate our method's robustness and scalability.
title Identifiable Object Representations under Spatial Ambiguities
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07806