MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamane, Taiga, Suzuki, Satoshi, Masumura, Ryo, Orihashi, Shota, Tanaka, Tomohiro, Ihori, Mana, Makishima, Naoki, Kawata, Naotaka
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916922926301184
author Yamane, Taiga
Suzuki, Satoshi
Masumura, Ryo
Orihashi, Shota
Tanaka, Tomohiro
Ihori, Mana
Makishima, Naoki
Kawata, Naotaka
author_facet Yamane, Taiga
Suzuki, Satoshi
Masumura, Ryo
Orihashi, Shota
Tanaka, Tomohiro
Ihori, Mana
Makishima, Naoki
Kawata, Naotaka
contents Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view (BEV) from multi-view images. In MVPD, end-to-end trainable deep learning methods have progressed greatly. However, they often struggle to detect pedestrians with consistently small or large scales in views or with vastly different scales between views. This is because they do not exploit multi-scale image features to generate the BEV feature and detect pedestrians. To overcome this problem, we propose a novel MVPD method, called Multi-Scale Multi-View Detection (MSMVD). MSMVD generates multi-scale BEV features by projecting multi-scale image features extracted from individual views into the BEV space, scale-by-scale. Each of these BEV features inherits the properties of its corresponding scale image features from multiple views. Therefore, these BEV features help the precise detection of pedestrians with consistently small or large scales in views. Then, MSMVD combines information at different scales of multiple views by processing the multi-scale BEV features using a feature pyramid network. This improves the detection of pedestrians with vastly different scales between views. Extensive experiments demonstrate that exploiting multi-scale image features via multi-scale BEV features greatly improves the detection performance, and MSMVD outperforms the previous highest MODA by $4.5$ points on the GMVD dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20447
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
Yamane, Taiga
Suzuki, Satoshi
Masumura, Ryo
Orihashi, Shota
Tanaka, Tomohiro
Ihori, Mana
Makishima, Naoki
Kawata, Naotaka
Computer Vision and Pattern Recognition
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view (BEV) from multi-view images. In MVPD, end-to-end trainable deep learning methods have progressed greatly. However, they often struggle to detect pedestrians with consistently small or large scales in views or with vastly different scales between views. This is because they do not exploit multi-scale image features to generate the BEV feature and detect pedestrians. To overcome this problem, we propose a novel MVPD method, called Multi-Scale Multi-View Detection (MSMVD). MSMVD generates multi-scale BEV features by projecting multi-scale image features extracted from individual views into the BEV space, scale-by-scale. Each of these BEV features inherits the properties of its corresponding scale image features from multiple views. Therefore, these BEV features help the precise detection of pedestrians with consistently small or large scales in views. Then, MSMVD combines information at different scales of multiple views by processing the multi-scale BEV features using a feature pyramid network. This improves the detection of pedestrians with vastly different scales between views. Extensive experiments demonstrate that exploiting multi-scale image features via multi-scale BEV features greatly improves the detection performance, and MSMVD outperforms the previous highest MODA by $4.5$ points on the GMVD dataset.
title MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.20447