Audio Geolocation: A Natural Sounds Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chasmai, Mustafa, Liu, Wuao, Maji, Subhransu, Van Horn, Grant
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915403863687168
author Chasmai, Mustafa
Liu, Wuao
Maji, Subhransu
Van Horn, Grant
author_facet Chasmai, Mustafa
Liu, Wuao
Maji, Subhransu
Van Horn, Grant
contents Can we determine someone's geographic location purely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? We tackle the challenge of global-scale audio geolocation, formalize the problem, and conduct an in-depth analysis with wildlife audio from the iNatSounds dataset. Adopting a vision-inspired approach, we convert audio recordings to spectrograms and benchmark existing image geolocation techniques. We hypothesize that species vocalizations offer strong geolocation cues due to their defined geographic ranges and propose an approach that integrates species range prediction with retrieval-based geolocation. We further evaluate whether geolocation improves when analyzing species-rich recordings or when aggregating across spatiotemporal neighborhoods. Finally, we introduce case studies from movies to explore multimodal geolocation using both audio and visual content. Our work highlights the advantages of integrating audio and visual cues, and sets the stage for future research in audio geolocation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio Geolocation: A Natural Sounds Benchmark
Chasmai, Mustafa
Liu, Wuao
Maji, Subhransu
Van Horn, Grant
Sound
Machine Learning
Audio and Speech Processing
Can we determine someone's geographic location purely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? We tackle the challenge of global-scale audio geolocation, formalize the problem, and conduct an in-depth analysis with wildlife audio from the iNatSounds dataset. Adopting a vision-inspired approach, we convert audio recordings to spectrograms and benchmark existing image geolocation techniques. We hypothesize that species vocalizations offer strong geolocation cues due to their defined geographic ranges and propose an approach that integrates species range prediction with retrieval-based geolocation. We further evaluate whether geolocation improves when analyzing species-rich recordings or when aggregating across spatiotemporal neighborhoods. Finally, we introduce case studies from movies to explore multimodal geolocation using both audio and visual content. Our work highlights the advantages of integrating audio and visual cues, and sets the stage for future research in audio geolocation.
title Audio Geolocation: A Natural Sounds Benchmark
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2505.18726