Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Bing, Jiang, Anbai, Zheng, Xinhu, Zhang, Wei-Qiang, Liu, Jia, Fan, Pingyi, Qian, Yanmin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918126411579392
author Han, Bing
Jiang, Anbai
Zheng, Xinhu
Zhang, Wei-Qiang
Liu, Jia
Fan, Pingyi
Qian, Yanmin
author_facet Han, Bing
Jiang, Anbai
Zheng, Xinhu
Zhang, Wei-Qiang
Liu, Jia
Fan, Pingyi
Qian, Yanmin
contents Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
Han, Bing
Jiang, Anbai
Zheng, Xinhu
Zhang, Wei-Qiang
Liu, Jia
Fan, Pingyi
Qian, Yanmin
Sound
Audio and Speech Processing
Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.
title Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.12230