Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head Tracking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Jisu, Mattingly, Alex, Lee, Eung-Joo, Riggan, Benjamin S.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908383325454336
author Kim, Jisu
Mattingly, Alex
Lee, Eung-Joo
Riggan, Benjamin S.
author_facet Kim, Jisu
Mattingly, Alex
Lee, Eung-Joo
Riggan, Benjamin S.
contents Head detection and tracking are essential for downstream tasks, but current methods often require large computational budgets, which increase latencies and ties up resources (e.g., processors, memory, and bandwidth). To address this, we propose a framework to enhance tiny head detection and tracking by optimizing the balance between performance and efficiency. Our framework integrates (1) a cross-domain detection loss, (2) a multi-scale module, and (3) a small receptive field detection mechanism. These innovations enhance detection by bridging the gap between large and small detectors, capturing high-frequency details at multiple scales during training, and using filters with small receptive fields to detect tiny heads. Evaluations on the CroHD and CrowdHuman datasets show improved Multiple Object Tracking Accuracy (MOTA) and mean Average Precision (mAP), demonstrating the effectiveness of our approach in crowded scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22677
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head Tracking
Kim, Jisu
Mattingly, Alex
Lee, Eung-Joo
Riggan, Benjamin S.
Computer Vision and Pattern Recognition
Head detection and tracking are essential for downstream tasks, but current methods often require large computational budgets, which increase latencies and ties up resources (e.g., processors, memory, and bandwidth). To address this, we propose a framework to enhance tiny head detection and tracking by optimizing the balance between performance and efficiency. Our framework integrates (1) a cross-domain detection loss, (2) a multi-scale module, and (3) a small receptive field detection mechanism. These innovations enhance detection by bridging the gap between large and small detectors, capturing high-frequency details at multiple scales during training, and using filters with small receptive fields to detect tiny heads. Evaluations on the CroHD and CrowdHuman datasets show improved Multiple Object Tracking Accuracy (MOTA) and mean Average Precision (mAP), demonstrating the effectiveness of our approach in crowded scenes.
title Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22677