TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anand, Nishit, Seth, Ashish, Duraiswami, Ramani, Manocha, Dinesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917975482695680
author Anand, Nishit
Seth, Ashish
Duraiswami, Ramani
Manocha, Dinesh
author_facet Anand, Nishit
Seth, Ashish
Duraiswami, Ramani
Manocha, Dinesh
contents Audio-language models (ALMs) excel in zero-shot audio classification, a task where models classify previously unseen audio clips at test time by leveraging descriptive natural language prompts. We introduce TSPE (Task-Specific Prompt Ensemble), a simple, training-free hard prompting method that boosts ALEs' zero-shot performance by customizing prompts for diverse audio classification tasks. Rather than using generic template-based prompts like "Sound of a car" we generate context-rich prompts, such as "Sound of a car coming from a tunnel". Specifically, we leverage label information to identify suitable sound attributes, such as "loud" and "feeble", and appropriate sound sources, such as "tunnel" and "street" and incorporate this information into the prompts used by Audio-Language Models (ALMs) for audio classification. Further, to enhance audio-text alignment, we perform prompt ensemble across TSPE-generated task-specific prompts. When evaluated on 12 diverse audio classification datasets, TSPE improves performance across ALMs by showing an absolute improvement of 1.23-16.36% over vanilla zero-shot evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
Anand, Nishit
Seth, Ashish
Duraiswami, Ramani
Manocha, Dinesh
Sound
Artificial Intelligence
Computation and Language
Machine Learning
Audio and Speech Processing
Audio-language models (ALMs) excel in zero-shot audio classification, a task where models classify previously unseen audio clips at test time by leveraging descriptive natural language prompts. We introduce TSPE (Task-Specific Prompt Ensemble), a simple, training-free hard prompting method that boosts ALEs' zero-shot performance by customizing prompts for diverse audio classification tasks. Rather than using generic template-based prompts like "Sound of a car" we generate context-rich prompts, such as "Sound of a car coming from a tunnel". Specifically, we leverage label information to identify suitable sound attributes, such as "loud" and "feeble", and appropriate sound sources, such as "tunnel" and "street" and incorporate this information into the prompts used by Audio-Language Models (ALMs) for audio classification. Further, to enhance audio-text alignment, we perform prompt ensemble across TSPE-generated task-specific prompts. When evaluated on 12 diverse audio classification datasets, TSPE improves performance across ALMs by showing an absolute improvement of 1.23-16.36% over vanilla zero-shot evaluation.
title TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
topic Sound
Artificial Intelligence
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2501.00398