Rethinking Global Context in Crowd Counting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Guolei, Liu, Yun, Probst, Thomas, Paudel, Danda Pani, Popovic, Nikola, Van Gool, Luc
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911813543657472
author Sun, Guolei
Liu, Yun
Probst, Thomas
Paudel, Danda Pani
Popovic, Nikola
Van Gool, Luc
author_facet Sun, Guolei
Liu, Yun
Probst, Thomas
Paudel, Danda Pani
Popovic, Nikola
Van Gool, Luc
contents This paper investigates the role of global context for crowd counting. Specifically, a pure transformer is used to extract features with global information from overlapping image patches. Inspired by classification, we add a context token to the input sequence, to facilitate information exchange with tokens corresponding to image patches throughout transformer layers. Due to the fact that transformers do not explicitly model the tried-and-true channel-wise interactions, we propose a token-attention module (TAM) to recalibrate encoded features through channel-wise attention informed by the context token. Beyond that, it is adopted to predict the total person count of the image through regression-token module (RTM). Extensive experiments on various datasets, including ShanghaiTech, UCF-QNRF, JHU-CROWD++ and NWPU, demonstrate that the proposed context extraction techniques can significantly improve the performance over the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2105_10926
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Rethinking Global Context in Crowd Counting
Sun, Guolei
Liu, Yun
Probst, Thomas
Paudel, Danda Pani
Popovic, Nikola
Van Gool, Luc
Computer Vision and Pattern Recognition
This paper investigates the role of global context for crowd counting. Specifically, a pure transformer is used to extract features with global information from overlapping image patches. Inspired by classification, we add a context token to the input sequence, to facilitate information exchange with tokens corresponding to image patches throughout transformer layers. Due to the fact that transformers do not explicitly model the tried-and-true channel-wise interactions, we propose a token-attention module (TAM) to recalibrate encoded features through channel-wise attention informed by the context token. Beyond that, it is adopted to predict the total person count of the image through regression-token module (RTM). Extensive experiments on various datasets, including ShanghaiTech, UCF-QNRF, JHU-CROWD++ and NWPU, demonstrate that the proposed context extraction techniques can significantly improve the performance over the baselines.
title Rethinking Global Context in Crowd Counting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2105.10926