Text-promptable Object Counting via Quantity Awareness Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Miaojing, Zhang, Xiaowen, Yue, Zijie, Luo, Yong, Zhao, Cairong, Li, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912472971083776
author Shi, Miaojing
Zhang, Xiaowen
Yue, Zijie
Luo, Yong
Zhao, Cairong
Li, Li
author_facet Shi, Miaojing
Zhang, Xiaowen
Yue, Zijie
Luo, Yong
Zhao, Cairong
Li, Li
contents Recent advances in large vision-language models (VLMs) have shown remarkable progress in solving the text-promptable object counting problem. Representative methods typically specify text prompts with object category information in images. This however is insufficient for training the model to accurately distinguish the number of objects in the counting task. To this end, we propose QUANet, which introduces novel quantity-oriented text prompts with a vision-text quantity alignment loss to enhance the model's quantity awareness. Moreover, we propose a dual-stream adaptive counting decoder consisting of a Transformer stream, a CNN stream, and a number of Transformer-to-CNN enhancement adapters (T2C-adapters) for density map prediction. The T2C-adapters facilitate the effective knowledge communication and aggregation between the Transformer and CNN streams. A cross-stream quantity ranking loss is proposed in the end to optimize the ranking orders of predictions from the two streams. Extensive experiments on standard benchmarks such as FSC-147, CARPK, PUCPR+, and ShanghaiTech demonstrate our model's strong generalizability for zero-shot class-agnostic counting. Code is available at https://github.com/viscom-tongji/QUANet
format Preprint
id arxiv_https___arxiv_org_abs_2507_06679
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-promptable Object Counting via Quantity Awareness Enhancement
Shi, Miaojing
Zhang, Xiaowen
Yue, Zijie
Luo, Yong
Zhao, Cairong
Li, Li
Computer Vision and Pattern Recognition
Recent advances in large vision-language models (VLMs) have shown remarkable progress in solving the text-promptable object counting problem. Representative methods typically specify text prompts with object category information in images. This however is insufficient for training the model to accurately distinguish the number of objects in the counting task. To this end, we propose QUANet, which introduces novel quantity-oriented text prompts with a vision-text quantity alignment loss to enhance the model's quantity awareness. Moreover, we propose a dual-stream adaptive counting decoder consisting of a Transformer stream, a CNN stream, and a number of Transformer-to-CNN enhancement adapters (T2C-adapters) for density map prediction. The T2C-adapters facilitate the effective knowledge communication and aggregation between the Transformer and CNN streams. A cross-stream quantity ranking loss is proposed in the end to optimize the ranking orders of predictions from the two streams. Extensive experiments on standard benchmarks such as FSC-147, CARPK, PUCPR+, and ShanghaiTech demonstrate our model's strong generalizability for zero-shot class-agnostic counting. Code is available at https://github.com/viscom-tongji/QUANet
title Text-promptable Object Counting via Quantity Awareness Enhancement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.06679