Acoustic Simulation Framework for Multi-channel Replay Speech Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neri, Michael, Virtanen, Tuomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916062602199040
author Neri, Michael
Virtanen, Tuomas
author_facet Neri, Michael
Virtanen, Tuomas
contents Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection robustness, existing datasets and methods predominantly rely on single-channel recordings. Moreover, previous studies highlighted that generalization of this attack to new environments is challenging, requiring new methods for generating data encompassing various acoustic conditions. Hence, in this work we introduce an acoustic simulation framework designed to simulate multi-channel replay speech configurations using publicly available resources. Using the framework, we train the state-of-the-art multi-channel replay detector M-ALRAD and evaluate its generalisation on the ReMASC real-recording corpus without any real training data. To improve the exploitation of spatial information, we extend M-ALRAD with inter-channel phase difference features computed for adjacent microphone pairs, augmenting the beamformed representation with directional cues. Synthetic datasets will be available upon acceptance of the paper.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Acoustic Simulation Framework for Multi-channel Replay Speech Detection
Neri, Michael
Virtanen, Tuomas
Audio and Speech Processing
Cryptography and Security
Sound
Signal Processing
Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection robustness, existing datasets and methods predominantly rely on single-channel recordings. Moreover, previous studies highlighted that generalization of this attack to new environments is challenging, requiring new methods for generating data encompassing various acoustic conditions. Hence, in this work we introduce an acoustic simulation framework designed to simulate multi-channel replay speech configurations using publicly available resources. Using the framework, we train the state-of-the-art multi-channel replay detector M-ALRAD and evaluate its generalisation on the ReMASC real-recording corpus without any real training data. To improve the exploitation of spatial information, we extend M-ALRAD with inter-channel phase difference features computed for adjacent microphone pairs, augmenting the beamformed representation with directional cues. Synthetic datasets will be available upon acceptance of the paper.
title Acoustic Simulation Framework for Multi-channel Replay Speech Detection
topic Audio and Speech Processing
Cryptography and Security
Sound
Signal Processing
url https://arxiv.org/abs/2509.14789