Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: N, Rishith Sadashiv T, Bedge, Abhishek, Bore, Saisha Suresh, Mishra, Jagabandhu, Bhattacharjee, Mrinmoy, Prasanna, S R Mahadeva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914006714810368
author N, Rishith Sadashiv T
Bedge, Abhishek
Bore, Saisha Suresh
Mishra, Jagabandhu
Bhattacharjee, Mrinmoy
Prasanna, S R Mahadeva
author_facet N, Rishith Sadashiv T
Bedge, Abhishek
Bore, Saisha Suresh
Mishra, Jagabandhu
Bhattacharjee, Mrinmoy
Prasanna, S R Mahadeva
contents Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to address domain generalization issue by proposing a novel speech representation using self-supervised (SSL) speech embeddings and the Modulation Spectrogram (MS) feature. A fusion strategy is used to combine both speech representations to introduce a new front-end for the classification task. The proposed SSL+MS fusion representation is passed to the AASIST back-end network. Experiments are conducted on monolingual and multilingual fake speech datasets to evaluate the efficacy of the proposed model architecture in cross-dataset and multilingual cases. The proposed model achieves a relative performance improvement of 37% and 20% on the ASVspoof 2019 and MLAAD datasets, respectively, in in-domain settings compared to the baseline. In the out-of-domain scenario, the model trained on ASVspoof 2019 shows a 36% relative improvement when evaluated on the MLAAD dataset. Across all evaluated languages, the proposed model consistently outperforms the baseline, indicating enhanced domain generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
N, Rishith Sadashiv T
Bedge, Abhishek
Bore, Saisha Suresh
Mishra, Jagabandhu
Bhattacharjee, Mrinmoy
Prasanna, S R Mahadeva
Audio and Speech Processing
Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to address domain generalization issue by proposing a novel speech representation using self-supervised (SSL) speech embeddings and the Modulation Spectrogram (MS) feature. A fusion strategy is used to combine both speech representations to introduce a new front-end for the classification task. The proposed SSL+MS fusion representation is passed to the AASIST back-end network. Experiments are conducted on monolingual and multilingual fake speech datasets to evaluate the efficacy of the proposed model architecture in cross-dataset and multilingual cases. The proposed model achieves a relative performance improvement of 37% and 20% on the ASVspoof 2019 and MLAAD datasets, respectively, in in-domain settings compared to the baseline. In the out-of-domain scenario, the model trained on ASVspoof 2019 shows a 36% relative improvement when evaluated on the MLAAD dataset. Across all evaluated languages, the proposed model consistently outperforms the baseline, indicating enhanced domain generalization.
title Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.01034