Leveraging Audio Representations for Vibration-Based Crowd Monitoring in Stadiums

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Yen Cheng, Codling, Jesse, Dong, Yiwen, Zhang, Jiale, Chen, Jiasi, Noh, Hae Young, Zhang, Pei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909547870814208
author Chang, Yen Cheng
Codling, Jesse
Dong, Yiwen
Zhang, Jiale
Chen, Jiasi
Noh, Hae Young
Zhang, Pei
author_facet Chang, Yen Cheng
Codling, Jesse
Dong, Yiwen
Zhang, Jiale
Chen, Jiasi
Noh, Hae Young
Zhang, Pei
contents Crowd monitoring in sports stadiums is important to enhance public safety and improve the audience experience. Existing approaches mainly rely on cameras and microphones, which can cause significant disturbances and often raise privacy concerns. In this paper, we sense floor vibration, which provides a less disruptive and more non-intrusive way of crowd sensing, to predict crowd behavior. However, since the vibration-based crowd monitoring approach is newly developed, one main challenge is the lack of training data due to sports stadiums being large public spaces with complex physical activities. In this paper, we present ViLA (Vibration Leverage Audio), a vibration-based method that reduces the dependency on labeled data by pre-training with unlabeled cross-modality data. ViLA is first pre-trained on audio data in an unsupervised manner and then fine-tuned with a minimal amount of in-domain vibration data. By leveraging publicly available audio datasets, ViLA learns the wave behaviors from audio and then adapts the representation to vibration, reducing the reliance on domain-specific vibration data. Our real-world experiments demonstrate that pre-training the vibration model using publicly available audio data (YouTube8M) achieved up to a 5.8x error reduction compared to the model without audio pre-training.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Audio Representations for Vibration-Based Crowd Monitoring in Stadiums
Chang, Yen Cheng
Codling, Jesse
Dong, Yiwen
Zhang, Jiale
Chen, Jiasi
Noh, Hae Young
Zhang, Pei
Sound
Computer Vision and Pattern Recognition
Crowd monitoring in sports stadiums is important to enhance public safety and improve the audience experience. Existing approaches mainly rely on cameras and microphones, which can cause significant disturbances and often raise privacy concerns. In this paper, we sense floor vibration, which provides a less disruptive and more non-intrusive way of crowd sensing, to predict crowd behavior. However, since the vibration-based crowd monitoring approach is newly developed, one main challenge is the lack of training data due to sports stadiums being large public spaces with complex physical activities. In this paper, we present ViLA (Vibration Leverage Audio), a vibration-based method that reduces the dependency on labeled data by pre-training with unlabeled cross-modality data. ViLA is first pre-trained on audio data in an unsupervised manner and then fine-tuned with a minimal amount of in-domain vibration data. By leveraging publicly available audio datasets, ViLA learns the wave behaviors from audio and then adapts the representation to vibration, reducing the reliance on domain-specific vibration data. Our real-world experiments demonstrate that pre-training the vibration model using publicly available audio data (YouTube8M) achieved up to a 5.8x error reduction compared to the model without audio pre-training.
title Leveraging Audio Representations for Vibration-Based Crowd Monitoring in Stadiums
topic Sound
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.17646