Listenable Maps for Zero-Shot Audio Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paissan, Francesco, Della Libera, Luca, Ravanelli, Mirco, Subakan, Cem
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909587110625280
author Paissan, Francesco
Della Libera, Luca
Ravanelli, Mirco
Subakan, Cem
author_facet Paissan, Francesco
Della Libera, Luca
Ravanelli, Mirco
Subakan, Cem
contents Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Audio Classifiers in the Zero-Shot context), which, to the best of our knowledge, is the first decoder-based post-hoc interpretation method for explaining the decisions of zero-shot audio classifiers. The proposed method utilizes a novel loss function that maximizes the faithfulness to the original similarity between a given text-and-audio pair. We provide an extensive evaluation using the Contrastive Language-Audio Pretraining (CLAP) model to showcase that our interpreter remains faithful to the decisions in a zero-shot classification context. Moreover, we qualitatively show that our method produces meaningful explanations that correlate well with different text prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Listenable Maps for Zero-Shot Audio Classifiers
Paissan, Francesco
Della Libera, Luca
Ravanelli, Mirco
Subakan, Cem
Sound
Machine Learning
Audio and Speech Processing
Signal Processing
Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Audio Classifiers in the Zero-Shot context), which, to the best of our knowledge, is the first decoder-based post-hoc interpretation method for explaining the decisions of zero-shot audio classifiers. The proposed method utilizes a novel loss function that maximizes the faithfulness to the original similarity between a given text-and-audio pair. We provide an extensive evaluation using the Contrastive Language-Audio Pretraining (CLAP) model to showcase that our interpreter remains faithful to the decisions in a zero-shot classification context. Moreover, we qualitatively show that our method produces meaningful explanations that correlate well with different text prompts.
title Listenable Maps for Zero-Shot Audio Classifiers
topic Sound
Machine Learning
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2405.17615