Grouped Discrete Representation for Object-Centric Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Rongzhen, Wang, Vivienne, Kannala, Juho, Pajarinen, Joni
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908636993814528
author Zhao, Rongzhen
Wang, Vivienne
Kannala, Juho
Pajarinen, Joni
author_facet Zhao, Rongzhen
Wang, Vivienne
Kannala, Juho
Pajarinen, Joni
contents Object-Centric Learning (OCL) aims to discover objects in images or videos by reconstructing the input. Representative methods achieve this by reconstructing the input as its Variational Autoencoder (VAE) discrete representations, which suppress (super-)pixel noise and enhance object separability. However, these methods treat features as indivisible units, overlooking their compositional attributes, and discretize features via scalar code indexes, losing attribute-level similarities and differences. We propose Grouped Discrete Representation (GDR) for OCL. For better generalization, features are decomposed into combinatorial attributes by organized channel grouping. For better convergence, features are quantized into discrete representations via tuple code indexes. Experiments demonstrate that GDR consistently improves both mainstream and state-of-the-art OCL methods across various datasets. Visualizations further highlight GDR's superior object separability and interpretability. The source code is available on https://github.com/Genera1Z/GroupedDiscreteRepresentation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02299
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grouped Discrete Representation for Object-Centric Learning
Zhao, Rongzhen
Wang, Vivienne
Kannala, Juho
Pajarinen, Joni
Computer Vision and Pattern Recognition
Machine Learning
Object-Centric Learning (OCL) aims to discover objects in images or videos by reconstructing the input. Representative methods achieve this by reconstructing the input as its Variational Autoencoder (VAE) discrete representations, which suppress (super-)pixel noise and enhance object separability. However, these methods treat features as indivisible units, overlooking their compositional attributes, and discretize features via scalar code indexes, losing attribute-level similarities and differences. We propose Grouped Discrete Representation (GDR) for OCL. For better generalization, features are decomposed into combinatorial attributes by organized channel grouping. For better convergence, features are quantized into discrete representations via tuple code indexes. Experiments demonstrate that GDR consistently improves both mainstream and state-of-the-art OCL methods across various datasets. Visualizations further highlight GDR's superior object separability and interpretability. The source code is available on https://github.com/Genera1Z/GroupedDiscreteRepresentation.
title Grouped Discrete Representation for Object-Centric Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.02299