The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Ruixing, Liu, Zihan, Sun, Leilei, Zhu, Tongyu, Lv, Weifeng
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915711431999488
author Zhang, Ruixing
Liu, Zihan
Sun, Leilei
Zhu, Tongyu
Lv, Weifeng
author_facet Zhang, Ruixing
Liu, Zihan
Sun, Leilei
Zhu, Tongyu
Lv, Weifeng
contents Geo-localization aims to infer the geographic origin of a given signal. In computer vision, geo-localization has served as a demanding benchmark for compositional reasoning and is relevant to public safety. In contrast, progress on audio geo-localization has been constrained by the lack of high-quality audio-location pairs. To address this gap, we introduce AGL1K, the first audio geo-localization benchmark for audio language models (ALMs), spanning 72 countries and territories. To extract reliably localizable samples from a crowd-sourced platform, we propose the Audio Localizability metric that quantifies the informativeness of each recording, yielding 1,444 curated audio clips. Evaluations on 16 ALMs show that ALMs have emerged with audio geo-localization capability. We find that closed-source models substantially outperform open-source models, and that linguistic clues often dominate as a scaffold for prediction. We further analyze ALMs' reasoning traces, regional bias, error causes, and the interpretability of the localizability metric. Overall, AGL1K establishes a benchmark for audio geo-localization and may advance ALMs with better geospatial reasoning capability.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03227
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
Zhang, Ruixing
Liu, Zihan
Sun, Leilei
Zhu, Tongyu
Lv, Weifeng
Sound
Artificial Intelligence
Geo-localization aims to infer the geographic origin of a given signal. In computer vision, geo-localization has served as a demanding benchmark for compositional reasoning and is relevant to public safety. In contrast, progress on audio geo-localization has been constrained by the lack of high-quality audio-location pairs. To address this gap, we introduce AGL1K, the first audio geo-localization benchmark for audio language models (ALMs), spanning 72 countries and territories. To extract reliably localizable samples from a crowd-sourced platform, we propose the Audio Localizability metric that quantifies the informativeness of each recording, yielding 1,444 curated audio clips. Evaluations on 16 ALMs show that ALMs have emerged with audio geo-localization capability. We find that closed-source models substantially outperform open-source models, and that linguistic clues often dominate as a scaffold for prediction. We further analyze ALMs' reasoning traces, regional bias, error causes, and the interpretability of the localizability metric. Overall, AGL1K establishes a benchmark for audio geo-localization and may advance ALMs with better geospatial reasoning capability.
title The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2601.03227