Human-like Object Grouping in Self-supervised Vision Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Adeli, Hossein, Ahn, Seoyoung, Luo, Andrew, Zhang, Mengmi, Kriegeskorte, Nikolaus, Zelinsky, Gregory
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911515620147200
author Adeli, Hossein
Ahn, Seoyoung
Luo, Andrew
Zhang, Mengmi
Kriegeskorte, Nikolaus
Zelinsky, Gregory
author_facet Adeli, Hossein
Ahn, Seoyoung
Luo, Andrew
Zhang, Mengmi
Kriegeskorte, Nikolaus
Zelinsky, Gregory
contents Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly understood. Here, we introduce a behavioral benchmark in which participants make same/different object judgments for dot pairs on naturalistic scenes, scaling up a classical psychophysics paradigm to over 1000 trials. We test a diverse set of vision models using a simple readout from their representations to predict subjects' reaction times. We observe a steady improvement across model generations, with both architecture and training objective contributing to alignment, and transformer-based models trained with the DINO self-supervised objective showing the strongest performance. To investigate the source of this improvement, we propose a novel metric to quantify the object-centric component of representations by measuring patch similarity within and between objects. Across models, stronger object-centric structure predicts human segmentation behavior more accurately. We further show that matching the Gram matrix of supervised transformer models, capturing similarity structure across image patches, with that of a self-supervised model through distillation improves their alignment with human behavior, converging with the prior finding that Gram anchoring improves DINOv3's feature quality. Together, these results demonstrate that self-supervised vision models capture object structure in a behaviorally human-like manner, and that Gram matrix structure plays a role in driving perceptual alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13994
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Human-like Object Grouping in Self-supervised Vision Transformers
Adeli, Hossein
Ahn, Seoyoung
Luo, Andrew
Zhang, Mengmi
Kriegeskorte, Nikolaus
Zelinsky, Gregory
Computer Vision and Pattern Recognition
Artificial Intelligence
Neurons and Cognition
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly understood. Here, we introduce a behavioral benchmark in which participants make same/different object judgments for dot pairs on naturalistic scenes, scaling up a classical psychophysics paradigm to over 1000 trials. We test a diverse set of vision models using a simple readout from their representations to predict subjects' reaction times. We observe a steady improvement across model generations, with both architecture and training objective contributing to alignment, and transformer-based models trained with the DINO self-supervised objective showing the strongest performance. To investigate the source of this improvement, we propose a novel metric to quantify the object-centric component of representations by measuring patch similarity within and between objects. Across models, stronger object-centric structure predicts human segmentation behavior more accurately. We further show that matching the Gram matrix of supervised transformer models, capturing similarity structure across image patches, with that of a self-supervised model through distillation improves their alignment with human behavior, converging with the prior finding that Gram anchoring improves DINOv3's feature quality. Together, these results demonstrate that self-supervised vision models capture object structure in a behaviorally human-like manner, and that Gram matrix structure plays a role in driving perceptual alignment.
title Human-like Object Grouping in Self-supervised Vision Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Neurons and Cognition
url https://arxiv.org/abs/2603.13994