Guardado en:
Detalles Bibliográficos
Autores principales: Hadjerci, Oussama, Letienne, Antoine, Hedjazi, Mohamed Abbas, Hafiane, Adel
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2508.19864
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916921244385280
author Hadjerci, Oussama
Letienne, Antoine
Hedjazi, Mohamed Abbas
Hafiane, Adel
author_facet Hadjerci, Oussama
Letienne, Antoine
Hedjazi, Mohamed Abbas
Hafiane, Adel
contents Self-supervised learning (SSL) has emerged as a powerful technique for learning visual representations. While recent SSL approaches achieve strong results in global image understanding, they are limited in capturing the structured representation in scenes. In this work, we propose a self-supervised approach that progressively builds structured visual representations by combining semantic grouping, instance level separation, and hierarchical structuring. Our approach, based on a novel ProtoScale module, captures visual elements across multiple spatial scales. Unlike common strategies like DINO that rely on random cropping and global embeddings, we preserve full scene context across augmented views to improve performance in dense prediction tasks. We validate our method on downstream object detection tasks using a combined subset of multiple datasets (COCO and UA-DETRAC). Experimental results show that our method learns object centric representations that enhance supervised object detection and outperform the state-of-the-art methods, even when trained with limited annotated data and fewer fine-tuning epochs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-supervised structured object representation learning
Hadjerci, Oussama
Letienne, Antoine
Hedjazi, Mohamed Abbas
Hafiane, Adel
Computer Vision and Pattern Recognition
Self-supervised learning (SSL) has emerged as a powerful technique for learning visual representations. While recent SSL approaches achieve strong results in global image understanding, they are limited in capturing the structured representation in scenes. In this work, we propose a self-supervised approach that progressively builds structured visual representations by combining semantic grouping, instance level separation, and hierarchical structuring. Our approach, based on a novel ProtoScale module, captures visual elements across multiple spatial scales. Unlike common strategies like DINO that rely on random cropping and global embeddings, we preserve full scene context across augmented views to improve performance in dense prediction tasks. We validate our method on downstream object detection tasks using a combined subset of multiple datasets (COCO and UA-DETRAC). Experimental results show that our method learns object centric representations that enhance supervised object detection and outperform the state-of-the-art methods, even when trained with limited annotated data and fewer fine-tuning epochs.
title Self-supervised structured object representation learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.19864