Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Loakman, Tyler, James, Joseph, Lin, Chenghua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909908226539520
author Loakman, Tyler
James, Joseph
Lin, Chenghua
author_facet Loakman, Tyler
James, Joseph
Lin, Chenghua
contents With the rise of Large Language Models (LLMs) and their vision-enabled counterparts (VLMs), numerous works have investigated their capabilities in tasks that fuse the modalities of vision and language. In this work, we benchmark the extent to which VLMs are able to act as highly-trained phoneticians, interpreting spectrograms and waveforms of speech. To do this, we synthesise a novel dataset containing 4k+ English words spoken in isolation alongside stylistically consistent spectrogram and waveform figures. We test the ability of VLMs to understand these representations of speech through a multiple-choice task whereby models must predict the correct phonemic or graphemic transcription of a spoken word when presented amongst 3 distractor transcriptions that have been selected based on their phonemic edit distance to the ground truth. We observe that both zero-shot and finetuned models rarely perform above chance, demonstrating the requirement for specific parametric knowledge of how to interpret such figures, rather than paired samples alone.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
Loakman, Tyler
James, Joseph
Lin, Chenghua
Computation and Language
With the rise of Large Language Models (LLMs) and their vision-enabled counterparts (VLMs), numerous works have investigated their capabilities in tasks that fuse the modalities of vision and language. In this work, we benchmark the extent to which VLMs are able to act as highly-trained phoneticians, interpreting spectrograms and waveforms of speech. To do this, we synthesise a novel dataset containing 4k+ English words spoken in isolation alongside stylistically consistent spectrogram and waveform figures. We test the ability of VLMs to understand these representations of speech through a multiple-choice task whereby models must predict the correct phonemic or graphemic transcription of a spoken word when presented amongst 3 distractor transcriptions that have been selected based on their phonemic edit distance to the ground truth. We observe that both zero-shot and finetuned models rarely perform above chance, demonstrating the requirement for specific parametric knowledge of how to interpret such figures, rather than paired samples alone.
title Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
topic Computation and Language
url https://arxiv.org/abs/2511.13225