Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khoury, Karim El, Zanella, Maxime, Godelaine, Tiffanie, De Vleeschouwer, Christophe, Macq, Benoit
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918318512799744
author Khoury, Karim El
Zanella, Maxime
Godelaine, Tiffanie
De Vleeschouwer, Christophe
Macq, Benoit
author_facet Khoury, Karim El
Zanella, Maxime
Godelaine, Tiffanie
De Vleeschouwer, Christophe
Macq, Benoit
contents Audio-language models have recently demonstrated strong zero-shot capabilities by leveraging natural-language supervision to classify audio events without labeled training data. Yet, their performance is highly sensitive to the wording of text prompts, with small variations leading to large fluctuations in accuracy. Prior work has mitigated this issue through prompt learning or prompt ensembling. However, these strategies either require annotated data or fail to account for the fact that some prompts may negatively impact performance. In this work, we present an entropy-guided prompt weighting approach that aims to find a robust combination of prompt contributions to maximize prediction confidence. To this end, we formulate a tailored objective function that minimizes prediction entropy to yield new prompt weights, utilizing low-entropy as a proxy for high confidence. Our approach can be applied to individual samples or a batch of audio samples, requiring no additional labels and incurring negligible computational overhead. Experiments on five audio classification datasets covering environmental, urban, and vocal sounds, demonstrate consistent gains compared to classical prompt ensembling methods in a zero-shot setting, with accuracy improvements 5-times larger across the whole benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05011
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification
Khoury, Karim El
Zanella, Maxime
Godelaine, Tiffanie
De Vleeschouwer, Christophe
Macq, Benoit
Sound
Machine Learning
Audio-language models have recently demonstrated strong zero-shot capabilities by leveraging natural-language supervision to classify audio events without labeled training data. Yet, their performance is highly sensitive to the wording of text prompts, with small variations leading to large fluctuations in accuracy. Prior work has mitigated this issue through prompt learning or prompt ensembling. However, these strategies either require annotated data or fail to account for the fact that some prompts may negatively impact performance. In this work, we present an entropy-guided prompt weighting approach that aims to find a robust combination of prompt contributions to maximize prediction confidence. To this end, we formulate a tailored objective function that minimizes prediction entropy to yield new prompt weights, utilizing low-entropy as a proxy for high confidence. Our approach can be applied to individual samples or a batch of audio samples, requiring no additional labels and incurring negligible computational overhead. Experiments on five audio classification datasets covering environmental, urban, and vocal sounds, demonstrate consistent gains compared to classical prompt ensembling methods in a zero-shot setting, with accuracy improvements 5-times larger across the whole benchmark.
title Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification
topic Sound
Machine Learning
url https://arxiv.org/abs/2601.05011