CountingDINO: A Training-free Pipeline for Class-Agnostic Counting using Unsupervised Backbones

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pacini, Giacomo, Bianchi, Lorenzo, Ciampi, Luca, Messina, Nicola, Amato, Giuseppe, Falchi, Fabrizio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909897216491520
author Pacini, Giacomo
Bianchi, Lorenzo
Ciampi, Luca
Messina, Nicola
Amato, Giuseppe
Falchi, Fabrizio
author_facet Pacini, Giacomo
Bianchi, Lorenzo
Ciampi, Luca
Messina, Nicola
Amato, Giuseppe
Falchi, Fabrizio
contents Class-agnostic counting (CAC) aims to estimate the number of objects in images without being restricted to predefined categories. However, while current exemplar-based CAC methods offer flexibility at inference time, they still rely heavily on labeled data for training, which limits scalability and generalization to many downstream use cases. In this paper, we introduce CountingDINO, the first training-free exemplar-based CAC framework that exploits a fully unsupervised feature extractor. Specifically, our approach employs self-supervised vision-only backbones to extract object-aware features, and it eliminates the need for annotated data throughout the entire proposed pipeline. At inference time, we extract latent object prototypes via ROI-Align from DINO features and use them as convolutional kernels to generate similarity maps. These are then transformed into density maps through a simple yet effective normalization scheme. We evaluate our approach on the FSC-147 benchmark, where we consistently outperform a baseline based on an SOTA unsupervised object detector under the same label- and training-free setting. Additionally, we achieve competitive results -- and in some cases surpass -- training-free methods that rely on supervised backbones, non-training-free unsupervised methods, as well as several fully supervised SOTA approaches. This demonstrates that label- and training-free CAC can be both scalable and effective. Code: https://lorebianchi98.github.io/CountingDINO/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16570
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CountingDINO: A Training-free Pipeline for Class-Agnostic Counting using Unsupervised Backbones
Pacini, Giacomo
Bianchi, Lorenzo
Ciampi, Luca
Messina, Nicola
Amato, Giuseppe
Falchi, Fabrizio
Computer Vision and Pattern Recognition
Class-agnostic counting (CAC) aims to estimate the number of objects in images without being restricted to predefined categories. However, while current exemplar-based CAC methods offer flexibility at inference time, they still rely heavily on labeled data for training, which limits scalability and generalization to many downstream use cases. In this paper, we introduce CountingDINO, the first training-free exemplar-based CAC framework that exploits a fully unsupervised feature extractor. Specifically, our approach employs self-supervised vision-only backbones to extract object-aware features, and it eliminates the need for annotated data throughout the entire proposed pipeline. At inference time, we extract latent object prototypes via ROI-Align from DINO features and use them as convolutional kernels to generate similarity maps. These are then transformed into density maps through a simple yet effective normalization scheme. We evaluate our approach on the FSC-147 benchmark, where we consistently outperform a baseline based on an SOTA unsupervised object detector under the same label- and training-free setting. Additionally, we achieve competitive results -- and in some cases surpass -- training-free methods that rely on supervised backbones, non-training-free unsupervised methods, as well as several fully supervised SOTA approaches. This demonstrates that label- and training-free CAC can be both scalable and effective. Code: https://lorebianchi98.github.io/CountingDINO/.
title CountingDINO: A Training-free Pipeline for Class-Agnostic Counting using Unsupervised Backbones
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.16570