Robustifying automatic speech recognition by extracting slowly varying features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pizarro, Matías, Kolossa, Dorothea, Fischer, Asja
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909378371649536
author Pizarro, Matías
Kolossa, Dorothea
Fischer, Asja
author_facet Pizarro, Matías
Kolossa, Dorothea
Fischer, Asja
contents In the past few years, it has been shown that deep learning systems are highly vulnerable under attacks with adversarial examples. Neural-network-based automatic speech recognition (ASR) systems are no exception. Targeted and untargeted attacks can modify an audio input signal in such a way that humans still recognise the same words, while ASR systems are steered to predict a different transcription. In this paper, we propose a defense mechanism against targeted adversarial attacks consisting in removing fast-changing features from the audio signals, either by applying slow feature analysis, a low-pass filter, or both, before feeding the input to the ASR system. We perform an empirical analysis of hybrid ASR models trained on data pre-processed in such a way. While the resulting models perform quite well on benign data, they are significantly more robust against targeted adversarial attacks: Our final, proposed model shows a performance on clean data similar to the baseline model, while being more than four times more robust.
format Preprint
id arxiv_https___arxiv_org_abs_2112_07400
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Robustifying automatic speech recognition by extracting slowly varying features
Pizarro, Matías
Kolossa, Dorothea
Fischer, Asja
Audio and Speech Processing
Machine Learning
Sound
In the past few years, it has been shown that deep learning systems are highly vulnerable under attacks with adversarial examples. Neural-network-based automatic speech recognition (ASR) systems are no exception. Targeted and untargeted attacks can modify an audio input signal in such a way that humans still recognise the same words, while ASR systems are steered to predict a different transcription. In this paper, we propose a defense mechanism against targeted adversarial attacks consisting in removing fast-changing features from the audio signals, either by applying slow feature analysis, a low-pass filter, or both, before feeding the input to the ASR system. We perform an empirical analysis of hybrid ASR models trained on data pre-processed in such a way. While the resulting models perform quite well on benign data, they are significantly more robust against targeted adversarial attacks: Our final, proposed model shows a performance on clean data similar to the baseline model, while being more than four times more robust.
title Robustifying automatic speech recognition by extracting slowly varying features
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2112.07400