Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saengthong, Phurich, Nishida, Tomoya, Dohi, Kota, Yamashita, Natsuo, Kawaguchi, Yohei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918177706868736
author Saengthong, Phurich
Nishida, Tomoya
Dohi, Kota
Yamashita, Natsuo
Kawaguchi, Yohei
author_facet Saengthong, Phurich
Nishida, Tomoya
Dohi, Kota
Yamashita, Natsuo
Kawaguchi, Yohei
contents Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect anomalies via nearest-neighbor search, but fine tuning on noisy machine sounds often acts like a denoising objective, suppressing noise and reducing generalization under mismatched mixtures or inconsistent labeling. Training-free systems with frozen self-supervised learning (SSL) encoders avoid this issue and show strong first-shot generalization, yet their performance drops when mixture embeddings deviate from clean-source embeddings. We propose to improve SSL backbones with a retain-not-denoise strategy that better preserves information from mixed sound sources. The approach combines a multi-label audio tagging loss with a mixture alignment loss that aligns student mixture embeddings to convex teacher embeddings of clean and noise inputs. Controlled experiments on stationary, non-stationary, and mismatched noise subsets demonstrate improved robustness under distribution shifts, narrowing the gap toward oracle mixture representations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25182
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
Saengthong, Phurich
Nishida, Tomoya
Dohi, Kota
Yamashita, Natsuo
Kawaguchi, Yohei
Audio and Speech Processing
Sound
Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect anomalies via nearest-neighbor search, but fine tuning on noisy machine sounds often acts like a denoising objective, suppressing noise and reducing generalization under mismatched mixtures or inconsistent labeling. Training-free systems with frozen self-supervised learning (SSL) encoders avoid this issue and show strong first-shot generalization, yet their performance drops when mixture embeddings deviate from clean-source embeddings. We propose to improve SSL backbones with a retain-not-denoise strategy that better preserves information from mixed sound sources. The approach combines a multi-label audio tagging loss with a mixture alignment loss that aligns student mixture embeddings to convex teacher embeddings of clean and noise inputs. Controlled experiments on stationary, non-stationary, and mismatched noise subsets demonstrate improved robustness under distribution shifts, narrowing the gap toward oracle mixture representations.
title Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.25182