A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Yuhong, Chen, Hongtao, Zhang, Qi, Chen, Jie, He, Zhaoxi, Liu, Mingzhe, Liao, Jianghai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915504869867520
author Feng, Yuhong
Chen, Hongtao
Zhang, Qi
Chen, Jie
He, Zhaoxi
Liu, Mingzhe
Liao, Jianghai
author_facet Feng, Yuhong
Chen, Hongtao
Zhang, Qi
Chen, Jie
He, Zhaoxi
Liu, Mingzhe
Liao, Jianghai
contents Accurate RGB-Thermal (RGB-T) crowd counting is crucial for public safety in challenging conditions. While recent Transformer-based methods excel at capturing global context, their inherent lack of spatial inductive bias causes attention to spread to irrelevant background regions, compromising crowd localization precision. Furthermore, effectively bridging the gap between these distinct modalities remains a major hurdle. To tackle this, we propose the Dual Modulation Framework, comprising two modules: Spatially Modulated Attention (SMA), which improves crowd localization by using a learnable Spatial Decay Mask to penalize attention between distant tokens and prevent focus from spreading to the background; and Adaptive Fusion Modulation (AFM), which implements a dynamic gating mechanism to prioritize the most reliable modality for adaptive cross-modal fusion. Extensive experiments on RGB-T crowd counting datasets demonstrate the superior performance of our method compared to previous works. Code available at https://github.com/Cht2924/RGBT-Crowd-Counting.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17079
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion
Feng, Yuhong
Chen, Hongtao
Zhang, Qi
Chen, Jie
He, Zhaoxi
Liu, Mingzhe
Liao, Jianghai
Computer Vision and Pattern Recognition
Accurate RGB-Thermal (RGB-T) crowd counting is crucial for public safety in challenging conditions. While recent Transformer-based methods excel at capturing global context, their inherent lack of spatial inductive bias causes attention to spread to irrelevant background regions, compromising crowd localization precision. Furthermore, effectively bridging the gap between these distinct modalities remains a major hurdle. To tackle this, we propose the Dual Modulation Framework, comprising two modules: Spatially Modulated Attention (SMA), which improves crowd localization by using a learnable Spatial Decay Mask to penalize attention between distant tokens and prevent focus from spreading to the background; and Adaptive Fusion Modulation (AFM), which implements a dynamic gating mechanism to prioritize the most reliable modality for adaptive cross-modal fusion. Extensive experiments on RGB-T crowd counting datasets demonstrate the superior performance of our method compared to previous works. Code available at https://github.com/Cht2924/RGBT-Crowd-Counting.
title A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17079