Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd Counting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Maochen, Li, Zekun, Zhang, Jian, Qi, Lei, Shi, Yinghuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913753899991040
author Yang, Maochen
Li, Zekun
Zhang, Jian
Qi, Lei
Shi, Yinghuan
author_facet Yang, Maochen
Li, Zekun
Zhang, Jian
Qi, Lei
Shi, Yinghuan
contents Semi-supervised crowd counting is crucial for addressing the high annotation costs of densely populated scenes. Although several methods based on pseudo-labeling have been proposed, it remains challenging to effectively and accurately utilize unlabeled data. In this paper, we propose a novel framework called Taste More Taste Better (TMTB), which emphasizes both data and model aspects. Firstly, we explore a data augmentation technique well-suited for the crowd counting task. By inpainting the background regions, this technique can effectively enhance data diversity while preserving the fidelity of the entire scenes. Secondly, we introduce the Visual State Space Model as backbone to capture the global context information from crowd scenes, which is crucial for extremely crowded, low-light, and adverse weather scenarios. In addition to the traditional regression head for exact prediction, we employ an Anti-Noise classification head to provide less exact but more accurate supervision, since the regression head is sensitive to noise in manual annotations. We conduct extensive experiments on four benchmark datasets and show that our method outperforms state-of-the-art methods by a large margin. Code is publicly available on https://github.com/syhien/taste_more_taste_better.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd Counting
Yang, Maochen
Li, Zekun
Zhang, Jian
Qi, Lei
Shi, Yinghuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Semi-supervised crowd counting is crucial for addressing the high annotation costs of densely populated scenes. Although several methods based on pseudo-labeling have been proposed, it remains challenging to effectively and accurately utilize unlabeled data. In this paper, we propose a novel framework called Taste More Taste Better (TMTB), which emphasizes both data and model aspects. Firstly, we explore a data augmentation technique well-suited for the crowd counting task. By inpainting the background regions, this technique can effectively enhance data diversity while preserving the fidelity of the entire scenes. Secondly, we introduce the Visual State Space Model as backbone to capture the global context information from crowd scenes, which is crucial for extremely crowded, low-light, and adverse weather scenarios. In addition to the traditional regression head for exact prediction, we employ an Anti-Noise classification head to provide less exact but more accurate supervision, since the regression head is sensitive to noise in manual annotations. We conduct extensive experiments on four benchmark datasets and show that our method outperforms state-of-the-art methods by a large margin. Code is publicly available on https://github.com/syhien/taste_more_taste_better.
title Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd Counting
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.17984