Expanding Zero-Shot Object Counting with Rich Prompts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Huilin, Li, Senyao, Yuan, Jingling, Yang, Zhengwei, Guo, Yu, Liu, Wenxuan, Zhong, Xian, He, Shengfeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912394462101504
author Zhu, Huilin
Li, Senyao
Yuan, Jingling
Yang, Zhengwei
Guo, Yu
Liu, Wenxuan
Zhong, Xian
He, Shengfeng
author_facet Zhu, Huilin
Li, Senyao
Yuan, Jingling
Yang, Zhengwei
Guo, Yu
Liu, Wenxuan
Zhong, Xian
He, Shengfeng
contents Expanding pre-trained zero-shot counting models to handle unseen categories requires more than simply adding new prompts, as this approach does not achieve the necessary alignment between text and visual features for accurate counting. We introduce RichCount, the first framework to address these limitations, employing a two-stage training strategy that enhances text encoding and strengthens the model's association with objects in images. RichCount improves zero-shot counting for unseen categories through two key objectives: (1) enriching text features with a feed-forward network and adapter trained on text-image similarity, thereby creating robust, aligned representations; and (2) applying this refined encoder to counting tasks, enabling effective generalization across diverse prompts and complex images. In this manner, RichCount goes beyond simple prompt expansion to establish meaningful feature alignment that supports accurate counting across novel categories. Extensive experiments on three benchmark datasets demonstrate the effectiveness of RichCount, achieving state-of-the-art performance in zero-shot counting and significantly enhancing generalization to unseen categories in open-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expanding Zero-Shot Object Counting with Rich Prompts
Zhu, Huilin
Li, Senyao
Yuan, Jingling
Yang, Zhengwei
Guo, Yu
Liu, Wenxuan
Zhong, Xian
He, Shengfeng
Computer Vision and Pattern Recognition
Expanding pre-trained zero-shot counting models to handle unseen categories requires more than simply adding new prompts, as this approach does not achieve the necessary alignment between text and visual features for accurate counting. We introduce RichCount, the first framework to address these limitations, employing a two-stage training strategy that enhances text encoding and strengthens the model's association with objects in images. RichCount improves zero-shot counting for unseen categories through two key objectives: (1) enriching text features with a feed-forward network and adapter trained on text-image similarity, thereby creating robust, aligned representations; and (2) applying this refined encoder to counting tasks, enabling effective generalization across diverse prompts and complex images. In this manner, RichCount goes beyond simple prompt expansion to establish meaningful feature alignment that supports accurate counting across novel categories. Extensive experiments on three benchmark datasets demonstrate the effectiveness of RichCount, achieving state-of-the-art performance in zero-shot counting and significantly enhancing generalization to unseen categories in open-world scenarios.
title Expanding Zero-Shot Object Counting with Rich Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.15398