Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Arcos-Holzinger, Sandra, Erfani, Sarah M., Bailey, James, Khudanpur, Sanjeev
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909012615757824
author Arcos-Holzinger, Sandra
Erfani, Sarah M.
Bailey, James
Khudanpur, Sanjeev
author_facet Arcos-Holzinger, Sandra
Erfani, Sarah M.
Bailey, James
Khudanpur, Sanjeev
contents Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation similarity or global dimensionality, offering limited visibility into local geometric changes. We ask: how do perturbations deform local geometry, and do these shifts track downstream automatic speech recognition (ASR) degradation? To address this, we present GRIDS, a framework using Local Intrinsic Dimensionality (LID) across layer-wise representations in WavLM and wav2vec 2.0. We find that LID increases for all low signal-to noise ratio (SNR) perturbations and diverges at high SNR: benign noise converges toward the clean profile, while adversarial inputs retain early-layer LID elevation. We show LID elevation co-occurs with increased WER, and that layer-wise LID features enable anomaly detection (AUROC 0.78-1.00), opening the door to transcript-free monitoring in S3Ms.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02715
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
Arcos-Holzinger, Sandra
Erfani, Sarah M.
Bailey, James
Khudanpur, Sanjeev
Audio and Speech Processing
Cryptography and Security
Machine Learning
Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation similarity or global dimensionality, offering limited visibility into local geometric changes. We ask: how do perturbations deform local geometry, and do these shifts track downstream automatic speech recognition (ASR) degradation? To address this, we present GRIDS, a framework using Local Intrinsic Dimensionality (LID) across layer-wise representations in WavLM and wav2vec 2.0. We find that LID increases for all low signal-to noise ratio (SNR) perturbations and diverges at high SNR: benign noise converges toward the clean profile, while adversarial inputs retain early-layer LID elevation. We show LID elevation co-occurs with increased WER, and that layer-wise LID features enable anomaly detection (AUROC 0.78-1.00), opening the door to transcript-free monitoring in S3Ms.
title Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
topic Audio and Speech Processing
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.02715