Vector-Quantized Vision Foundation Models for Object-Centric Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Rongzhen, Wang, Vivienne, Kannala, Juho, Pajarinen, Joni
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917068252643328
author Zhao, Rongzhen
Wang, Vivienne
Kannala, Juho
Pajarinen, Joni
author_facet Zhao, Rongzhen
Wang, Vivienne
Kannala, Juho
Pajarinen, Joni
contents Object-Centric Learning (OCL) aggregates image or video feature maps into object-level feature vectors, termed \textit{slots}. It's self-supervision of reconstructing the input from slots struggles with complex object textures, thus Vision Foundation Model (VFM) representations are used as the aggregation input and reconstruction target. Existing methods leverage VFM representations in diverse ways yet fail to fully exploit their potential. In response, we propose a unified architecture, Vector-Quantized VFMs for OCL (VQ-VFM-OCL, or VVO). The key to our unification is simply shared quantizing VFM representations in OCL aggregation and decoding. Experiments show that across different VFMs, aggregators and decoders, our VVO consistently outperforms baselines in object discovery and recognition, as well as downstream visual prediction and reasoning. We also mathematically analyze why VFM representations facilitate OCL aggregation and why their shared quantization as reconstruction targets strengthens OCL supervision. Our source code and model checkpoints are available on https://github.com/Genera1Z/VQ-VFM-OCL.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20263
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vector-Quantized Vision Foundation Models for Object-Centric Learning
Zhao, Rongzhen
Wang, Vivienne
Kannala, Juho
Pajarinen, Joni
Computer Vision and Pattern Recognition
Object-Centric Learning (OCL) aggregates image or video feature maps into object-level feature vectors, termed \textit{slots}. It's self-supervision of reconstructing the input from slots struggles with complex object textures, thus Vision Foundation Model (VFM) representations are used as the aggregation input and reconstruction target. Existing methods leverage VFM representations in diverse ways yet fail to fully exploit their potential. In response, we propose a unified architecture, Vector-Quantized VFMs for OCL (VQ-VFM-OCL, or VVO). The key to our unification is simply shared quantizing VFM representations in OCL aggregation and decoding. Experiments show that across different VFMs, aggregators and decoders, our VVO consistently outperforms baselines in object discovery and recognition, as well as downstream visual prediction and reasoning. We also mathematically analyze why VFM representations facilitate OCL aggregation and why their shared quantization as reconstruction targets strengthens OCL supervision. Our source code and model checkpoints are available on https://github.com/Genera1Z/VQ-VFM-OCL.
title Vector-Quantized Vision Foundation Models for Object-Centric Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.20263