Vision Language Models Are Few-Shot Audio Spectrogram Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dixit, Satvik, Heller, Laurie M., Donahue, Chris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912124181151744
author Dixit, Satvik
Heller, Laurie M.
Donahue, Chris
author_facet Dixit, Satvik
Heller, Laurie M.
Donahue, Chris
contents We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct VLMs to perform audio classification tasks in a few-shot setting by prompting them to classify a spectrogram image given example spectrogram images of each class. By carefully designing the spectrogram image representation and selecting good few-shot examples, we show that GPT-4o can achieve 59.00% cross-validated accuracy on the ESC-10 environmental sound classification dataset. Moreover, we demonstrate that VLMs currently outperform the only available commercial audio language model with audio understanding capabilities (Gemini-1.5) on the equivalent audio classification task (59.00% vs. 49.62%), and even perform slightly better than human experts on visual spectrogram classification (73.75% vs. 72.50% on first fold). We envision two potential use cases for these findings: (1) combining the spectrogram and language understanding capabilities of VLMs for audio caption augmentation, and (2) posing visual spectrogram classification as a challenge task for VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12058
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
Dixit, Satvik
Heller, Laurie M.
Donahue, Chris
Sound
Audio and Speech Processing
We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct VLMs to perform audio classification tasks in a few-shot setting by prompting them to classify a spectrogram image given example spectrogram images of each class. By carefully designing the spectrogram image representation and selecting good few-shot examples, we show that GPT-4o can achieve 59.00% cross-validated accuracy on the ESC-10 environmental sound classification dataset. Moreover, we demonstrate that VLMs currently outperform the only available commercial audio language model with audio understanding capabilities (Gemini-1.5) on the equivalent audio classification task (59.00% vs. 49.62%), and even perform slightly better than human experts on visual spectrogram classification (73.75% vs. 72.50% on first fold). We envision two potential use cases for these findings: (1) combining the spectrogram and language understanding capabilities of VLMs for audio caption augmentation, and (2) posing visual spectrogram classification as a challenge task for VLMs.
title Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.12058